Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,034 datasets
Remote Sensing Systems produced this dataset of microwave radiometer wind speed, rain, and cloud liquid water data collocated to ISS-RapidScat Level 2B wind vector cell locations. The collocated radiometer sources include DMSP SSM/I, SSMIS, WindSat, AMSR2, and GMI, providing multi-sourced observations for the same geophysical points. Data coverage is restricted to latitudes between approximately 61 degrees North and 61 degrees South due to the International Space Station's orbit.
DiffSpot is a dataset of rendered web interface screenshot pairs, each differing by a single mutated CSS property. The dataset was created by Tencent and last updated on the Hugging Face platform in May 2026. It serves as a probe for testing the fine-grained visual perception capabilities of vision-language models.
ProteinGym is the standard benchmark suite for evaluating protein fitness and mutation effect predictors, developed by the OATML / Marks lab and published at NeurIPS 2023 Datasets and Benchmarks. It provides a common, standardized task surface for comparing protein language models, inverse folding models, and supervised fitness predictors. The benchmark covers both zero-shot and supervised settings.
SpatialUncertain is a controlled 3D benchmark designed to evaluate Vision-Language Models (VLMs). It contains approximately 6,600 questions for occlusion and 3,700 for perspective ambiguity, testing whether models know when not to answer spatial reasoning questions. The dataset was created by Yuezhangjoslin and serves as a companion to the paper 'Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?', with a last recorded update in June 2026.
76.4% of British Columbia is restricted from solar farm development according to a study using the Analytic Hierarchy Process. The analysis categorized land suitability, finding 0.21% (2,026.91 km²) as suitable and 0.005% (51.50 km²) as very suitable for utility-scale solar farms. The study by Mackenzie Thomson provides a province-wide screening tool for photovoltaic installation potential.
KFUPM-JRCAI created this dataset of machine-generated Arabic text for research on detection methods. The data was generated using multiple methods and Large Language Models (LLMs). It supports the research paper 'Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation'.
A matrix representation of surface materials derived from RVBI aerial photographs. The dataset identifies vegetated and mineralized areas to support land planning and management. It was published by the Government and Municipalities of Québec and was last updated on April 22, 2026.
An overview of analyzed genomics datasets compiled by Linnea Blomberg. It includes tissue types, pairwise comparisons, Gene Expression Omnibus (GEO) accession numbers, data types, data generation methods, and sample counts. The dataset is small at 13.5 KB and was last updated on May 15, 2026.
The dataset contains real-time observations from a global array of over 3,000 autonomous profiling floats. Each float cycles approximately every 10 days, recording temperature and salinity measurements from the upper 2000 meters of the ice-free oceans, along with mid-depth current data. Argo Australia contributes specific observations from the oceans surrounding Australia, with data mirrored and updated regularly.
Ziluo Fang created a dataset comparing the results of trigonometric function generation and quadratic polynomial generation. The dataset is stored as a 5.5 KB Excel file. It was last updated on May 15, 2026.
Uster Tester 5 (UT5) measurements for yarn spun from the premium Egyptian cotton variety "Giza 86". The dataset is published by author Rana Ahmed on figshare under a CC-BY-4.0 license. It was last updated on 2026-05-15.
Texas Department of Insurance records detail waiver requests from vision plans unable to meet state network adequacy standards. Each record tracks a request's status, insurer, county, specialty type, and hearing dates. The data is published by the Texas Department of Insurance on data.texas.gov.
The French government's 2014 Finance Bill (PLF 2014) for the general budget, categorized by destination and nature. The dataset is sourced from the official French open data portal, data.gouv.fr, and was uploaded to the platform by the account Data-Gouv-FR. The record was last updated on the platform in May 2026.
12,853 question-answer pairs generated from arXiv paper abstracts, covering computer science research across multiple subfields. The dataset was created by Navyasri12355 for instruction-tuning language models on scientific literature understanding tasks. It was last updated on the Hugging Face platform on 2026-05-29.
A framework for modelling shoreline response to clustered storm events focuses on two case study areas in southeast Australia: Adelaide metropolitan coast and Old Bar beach. The dataset likely contains coastal sediment compartment mapping, sub-surface sediment volume estimates, and event time series data. It is presented by the Australian Ocean Data Network as part of a Bushfire and Natural Hazard Cooperative Research Centre project.
Neurora's synthetic translation dataset was distilled from multiple open corpora using a teacher model for edge deployment. Designed for offline, low-latency machine translation on Android devices, the dataset was generated entirely using renewable energy. The dataset was last updated on May 25, 2026.
From the 17th to the 20th century, this dataset documents the territorial expansion of British India, detailing how provinces and districts were formed. It categorizes acquisitions by presidencies like Bengal, Madras, and Bombay, listing dates and methods of land acquisition. The dataset was curated by the India State Stories team using sources including Baden-Powell (1892) and India Administrative Atlas 1872-2001.
4,679,018 rows of knowledge graph triples for the telecommunications industry, compiled from Wikidata entities and SDO standards catalogs. The dataset was created by PleIAs and last updated on May 13, 2026. It includes 301,949 distinct subjects and is organized to highlight major telcos, equipment vendors, and 3GPP specifications.
Imagery from October-November 2020 comprises 2015 RGB images captured at 8300 feet. The orthorectified mosaic has a 6cm spatial resolution and a stated accuracy of 3 pixels at 68% confidence. Data was produced by the Bundaberg Regional Council GIS team.
227 RGB aerial images of Childers, captured on 2020-08-26 with a 300mm focal length sensor from 8300 feet. The imagery has been orthorectified into a digital orthophoto mosaic with a spatial resolution of 6 centimeters and a stated spatial accuracy of 3 pixels at 68% confidence. The data was produced by the Bundaberg Regional Council using Visionmap Lightspeed rectification processes.