Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
2018/19 records detail funerals arranged by the London Borough of Barnet for individuals with no next of kin or estate. The dataset includes personal identifiers, dates, addresses, and interactions with the Government Legal Department concerning bona vacantia (ownerless property). Columns suggest tracking of administrative processes for public health and estate management.
Over 30 columns detail each maintenance request, including its description, assigned title, request date, and completion date. The data originates from New York City's Asset Management Parks System and was last updated in March 2026. Each row represents a single work order, linkable to specific assets.
4006 raw agent trace files generated by the teich tool using the deepseek/deepseek-v4-pro model. The dataset is intended to prepare data for supervised fine-tuning in just a few lines of code. It was uploaded by ansulev and last updated on May 15, 2026.
NASA's Hurricane and Severe Storm Sentinel (HS3) CIMSS Brightness Temperature dataset provides infrared satellite imagery from GOES-15 and METEOSAT-10. The collection contains images captured at 15-minute intervals during the 2014 HS3 field campaign, spanning from August 14 to October 3. This dataset was created to study storm-scale processes and the role of the Saharan Air Layer in tropical cyclone formation.
A geospatial dataset from the USDA Forest Service, last updated March 13, 2026, representing watershed condition assessments on Forest Service lands. The data focuses on HUC12 watersheds with more than 5% USFS ownership and includes information on high-priority watersheds, designation rationales, and Watershed Restoration Action Plans. It is compiled from the NRM Watershed Condition Assessment and Tracking Tool (WCATT) application.
A dataset of 18.1 KB, last updated on 2026-04-20, examining psychological factors in using generative AI for health information. The dataset, authored by Nalae Hong and shared on figshare, likely contains survey responses measuring constructs like dependency, empowerment, and privacy risk. It is stored in the SAV format, commonly used for statistical software.
Olivia Plateau's dataset contains PLY and PTS files for avian palatine and pterygoid bones, used to study morphological disparity. The 319.5 MB collection includes an R script for running the main analyses from the associated paper. The data was last updated on April 27, 2026.
This dataset contains vertical scale measurements of near-field mixing driven by lee waves, dynamically constrained by inertial currents. Collected between 2020 and 2022, it supports research on ocean mixing processes and parameterizations.
Guardian FailCoT bundles three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model. The collection includes the UR5-Fail, RoboFail, and RoboVQA datasets. It was created by paulpacaud and last updated in May 2026.
The Canadian Hydrographic Service provides bathymetric data at 10-meter or 100-meter spatial resolution for non-navigational use. These products consolidate digital depth sources across Canadian jurisdiction, with horizontal referencing to WGS84 and vertical referencing to local Chart Datum. Data is organized into individual NONNA-10, NONNA-100, or packaged NONNA-P10 files with coverage varying by latitude.
Brisbane City Council's dataset maps areas subject to overland flow flooding within its local government area. It delineates high, medium, and low impact zones based on a 1% Annual Exceedance Probability flood extent from a 2017 mapping study. The dataset was created in June 2013 and is maintained by the council.
52,000 instruction-response pairs were translated from Russian to Bashkir for adapting large language models. The dataset was created by author metuKKhud and a small subset of about 10 examples was manually validated by a native speaker. It was last updated on May 20, 2026.
A 100,000-row subset of tokenized sequences for the EleutherAI/pythia-160m tokenizer, intended for language-model pre-pretraining experiments. The dataset was created by guox18 and last updated on May 20, 2026. Code and experiment scripts are available on GitHub.
The Ulysses/SWICS instrument measures solar wind ions from H through Fe across an energy per charge range from 0.16 to 59.6 keV/e with a time resolution of about 13 minutes. These data consist of 18 Matrix Rates (MR) as a function of energy per charge and time, representing specific elements and ionization states. The National Aeronautics and Space Administration provides the data, which can be processed with SAPRO software to convert count rates to physical units and kinetic parameters.
Global sea surface temperature measurements derived from the VIIRS sensor aboard the Suomi NPP satellite, launched on 28 October 2011. The data is produced by NOAA using the Advanced Clear-Sky Processor for Ocean (ACSPO) system and reported in 10-minute granules at a native resolution of approximately 0.75 km. This version 2.80 dataset includes algorithm improvements such as added thermal front layers and mitigated warm biases in high latitudes.
NOAA-20 satellite data provides sea surface temperature (SST) measurements at a native sensor resolution of approximately 0.75 km at nadir. The product is generated by NOAA's Advanced Clear-Sky Processor for Ocean (ACSPO) system, reported in 10-minute granules compliant with GHRSST standards, and validated against in situ data. Version 2.80 includes algorithm improvements such as added thermal front layers and mitigated warm biases in high latitudes.
Data.calgary.ca provides records of abandoned shopping carts on public property, collected via active 311 service requests. The dataset includes location coordinates, cart brand, condition, and associated bylaw infraction details. It was last updated in April 2026.
NOAA's GHRSST GOES-16 ABI L2P dataset provides sea surface temperature (SST) measurements for the Americas region from the GOES-East satellite. The data is derived from the Advanced Baseline Imager using a clear-sky processor and non-linear SST algorithm, producing 24 netCDF4 files per day with a total volume of 0.6GB. It is produced by the National Aeronautics and Space Administration and was last updated in March 2026.
Australian continental-scale pixel composites of surface reflectance for coastal and estuarine environments, corrected for tidal influences using a multi-resolution tidal model. The dataset includes high and low tide mosaics of the Australian coastline, generated by researchers from the Australian Ocean Data Network and published in 2018.
Continental-scale pixel composites of the Australian coastline are generated using a multi-resolution tidal model to account for dynamic water levels. The method employs a Voronoi mesh to capture spatial tidal variation and preserves spectral band relationships at each pixel. This dataset was created by researchers from the Australian Ocean Data Network and published in the journal Remote Sensing in 2018.