Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,606 datasets
Annual average employment statistics for Alberta from 2006 to 2017, disaggregated by immigrant status, sex, age group, and class of worker. The data is a customization of Statistics Canada information published by the Government of Alberta. It was last updated on April 17, 2026.
RUC: Real UI Clickboxes is a benchmark dataset for evaluating vision-language models and web agents on their ability to understand and resist deceptive user interfaces. The dataset accompanies an ACL submission and is released by the DUDE-Framework. It was last updated on Hugging Face on May 13, 2026.
Data used to build graphs for the article 'Adaptation of the spray generation function for spume droplets from a laboratory to a field environment'. The dataset, authored by Anna Zotova, is 2.0 MB in size and was last updated on May 8, 2026. It is available under a CC-BY-4.0 license.
A 5.2 KB dataset by Alejandro Hernández-Arango, last updated April 30, 2026, compares NLP-extracted and human-assigned IDSA/ATS clinical risk scores. Score distributions are presented under both labeling scenarios, along with the percentage of patients reclassified to a higher or lower risk stratum.
A 5.5 KB Excel file summarizing three speech emotion datasets used in machine learning experiments. The summary was authored by Mohammed Tawfik and last updated on May 7, 2026. It describes the roles of EmoDB and RAVDESS as federated training sources and CREMA-D for cross-corpus evaluation.
A 9.5 KB dataset listing bacterial genera that showed a significant increase in abundance within the intestinal microbiota of Nile tilapia fingerlings from a specific experimental group, C33D. The data is provided in an XLS file by author Mario Andrés Colorado Gómez and was last updated on May 7, -2026. The dataset is licensed under CC-BY-4.0.
5.5 KB Excel file lists bacterial genera showing significant abundance reduction in the intestinal microbiota of Nile tilapia fingerlings. The dataset, authored by Mario Andrés Colorado Gómez, was last updated on May 7, 2026, and is shared under a CC-BY-4.0 license on figshare.
Over 600 European companies' annual and sustainability reports from the STOXX Europe 600 index, covering a 10-year period. Kerstin Forster created this corpus for a study assessing corporate sustainability with large language models. Reports were sourced from company websites and reporting intermediaries to extract ESG indicators.
5.5 KB of per-class performance metrics for sleep quality classification and depressive sentiment analysis, authored by Akshi Kumar and last updated on May 7, 2026. The dataset is available in XLS format under a CC-BY-4.0 license on figshare.
Akshi Kumar published a dataset on figshare in May 2026 comparing model performance for two tasks. The dataset is a 5.5 KB Excel file containing results for sleep quality classification and depressive sentiment analysis. The specific data volume and column details are not provided in the metadata.
9.5 KB of tabular data from figshare compares patient baseline characteristics based on pet ownership. The dataset, authored by Hyeon Gyu Cho, is licensed under CC-BY-4.0 and was last updated on May 7, 2026. Its small size suggests a focused study or pilot analysis.
An inventory of fern and lycophyte families in the Sierra Madre Oriental mountain range in Mexico. The dataset records the number of species and genera per family, as well as endemic counts. It was authored by J. Daniel Tejero-Díez and last updated on 2026-05-07.
A botanical inventory lists families of ferns and lycophytes present in the Sierra Madre Oriental mountain range in Mexico. The dataset records the number of species and genera per family, including counts of endemic species. J. Daniel Tejero-Díez authored this 9.5 KB Excel file, last updated on May 7, -2026.
A 4.5-month-old boy's autopsy report details the first comprehensive pathological findings for CONDCA, a rare genetic neurodegeneration linked to AGTPBP1 gene mutation. The material includes histopathological descriptions of cerebellar atrophy, spinal cord degeneration, and systemic complications like thymic aplasia and adenovirus pneumonia. The case report and literature review were published by Karger via figshare in April 2026.
Situación de Caja para Regalías dataset provides an overview of approved resources and cash positions for Colombia's General Royalties System (SGR). It tracks investment budgets, collected revenues, cash balances, and disbursed payments for territorial entities and investment funds. The data is published by datos.gov.co and was last updated in March 2026.
Two roughly circular blue holes on the Pompey Reefs of the Great Barrier Reef, measuring 240-295 meters in diameter and 30-40 meters deep. The dataset, from the Australian Ocean Data Network, describes their distinct biological and sedimentary associations, steep inner slopes, and evidence suggesting they are collapsed dolines formed over multiple low sea-level periods. It was last updated on 2026-04-10.
PII Shield is a large-scale, multilingual dataset for training and evaluating Personally Identifiable Information detection models. Built by Auren Research, it combines real-world documents from diverse domains with high-quality span-level PII annotations produced by fastino/gliner2-privacy-filter-PII-multi. The dataset achieved the highest F1 score on the SPY benchmark among open-source PII detectors.
The XXXXXL Chain Of Thought dataset is hosted on HuggingFace by author 'wop' and was last updated on 2026-05-16. It is described as a reasoning style where an AI's internal thought process resembles a living stream of consciousness while performing low-level verification. The dataset appears to be designed for training or studying advanced reasoning models.
VIIRS/NPP Sea Ice Extent 6-Min L2 Swath 375m NRT data detects snow-covered sea ice using the Normalized Difference Snow Index (NDSI) algorithm, derived from radiance measurements by the Visible Infrared Imager Radiometer Suite aboard the Suomi NPP satellite. This near-real-time product provides high-resolution (375-meter) swath data every six minutes, reporting the location of sea ice cover. The methodology follows the approach established by the MODIS instrument, with detailed documentation including a Product User Guide and Algorithm Theoretical Basis Document available from NASA.
Forest Service FACTS data tracks vegetative manipulation activities designed to reduce wildland fire intensity and severity. This geospatial layer represents line features for hazardous fuel treatment reduction, including burning and mechanical methods. The data is managed by the USDA Forest Service's Natural Resource Manager and was last updated in March 2026.