Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,416 datasets
Logos v1.0 Corpus is a collection of source code programs written in the Logos programming language. The corpus includes cookbook recipes, algorithm examples, standard library code, and integration-test fixtures. It was created by byShammy and released on 2026-05-22.
NASA's New Horizons mission collected solar wind pick-up ion histogram data from the Solar Wind Around Pluto (SWAP) instrument. The dataset includes validated summary data and is associated with peer-reviewed publications from 2008 to 2017. Data files are available in BIN and HTML formats.
Momento is a benchmark dataset for evaluating LLM-based agents on persistent, tool-mediated task completion across multiple conversational sessions. The dataset, created by adrilmanurung, is grounded in a restaurant service domain and was last updated on May 30, 2026. Tasks require agents to recall past user preferences and resolve goals spanning multiple interactions.
A 5.5 KB Excel file containing nominal simulation parameters for a resilient distributed model predictive control (RDMPC) framework. The dataset, authored by Baheej Alghamdi and uploaded to figshare in April 2026, supports the coordination of five networked microgrids with demand response integration.
Experimental scenarios evaluate a resilient distributed model predictive control framework for networked microgrids. The 5.5 KB Excel file contains results from a five-microgrid benchmark under communication failures and topology changes. Author Baheej Alghamdi published the data on figshare in April 2026.
Execution semantics detail a resilient distributed model predictive control framework for networked microgrids. The 5.5 KB Excel file, authored by Baheej Alghamdi, was last updated in April 2026. It formalizes coordination mechanisms resilient to communication failures and integrates demand response models.
NASA's BOREAS AFM-11 team produced quality control and sampling analysis reports for data collected by Electra, Long-EZ, and Twin Otter aircraft. The dataset consists of archived PDF documents from the Boreal Ecosystem-Atmosphere Study. Platform update dates conflict, with one source listing a 1996 date and others showing a 2026 administrative update.
Parker Solar Probe SPAN-Ion instrument measurements of alpha particle differential energy flux, with data covering 8 deflector settings, 32 energies, and 8 angles. The dataset's energy range is log-spaced from 21.0 eV to 17.6 keV, and its field of view covers nearly the entire sky. It was produced by the National Aeronautics and Space Administration and last updated in March 2026.
420 training runs were conducted to develop a multi-agent reinforcement learning model for collaborative energy optimization in power routers. The model, proposed by Junyan Lyu, was last updated in April 2026 and is documented in a 9.5 KB XLS file. It demonstrates performance metrics including a stable average reward and reduced operational costs.
150 human-written and 150 AI-generated abstracts from high-impact linguistics and computer science journals were analyzed for readability and writing style. The dataset contains quantitative metrics from a Readability Scoring System and expert evaluations, compiled by Yumei Zou and last updated in April 2026. Analysis was performed using SPSS 27 with non-parametric statistical methods.
2838 complete and 269 partially complete 1:20,000 National Topographic System (NTS) grid blocks cover the province of Alberta. The grid was developed for the Provincial Digital Base Mapping Project, with each polygon designated by a structured alphanumeric code (e.g., 83H08NE). It is produced for the Government of Alberta and made available to the public.
Neiva Municipality in Colombia provides a list of miners and barequeros (artisanal miners). The dataset includes columns for MATERIAL DE EXPLOTACIΓN, SEXO, FUENTE HIDRICA DE LABOR, and NOMBRES. It was published on the Socrata platform via datos.gov.co and was last updated on 2026-05-18.
An Annotation for Transparent Inquiry (ATI) project analyzing perceptions of fairness in criminal justice decisions. The qualitative component focuses on a debate between ProPublica and Northpointe regarding the COMPAS risk assessment tool, while the quantitative component uses a dataset of over 7,000 individuals compiled by ProPublica. The data was gathered via Freedom of Information Act requests and analyzed using critical discourse analysis and SPSS.
From 2019 to 2021, Dana Mahr conducted 20 qualitative interviews with users of the PatientsLikeMe health platform. The interviews, lasting 60-90 minutes, focused on experiences with data sharing, privacy perceptions, and data commodification. The data was pseudonymized and analyzed using a grounded theory approach.
GLM-5.1-Thinking Distilled Dataset contains 5,000 unique examples of questions paired with detailed multi-step reasoning traces and final responses. Authored by gss1147 and hosted on Hugging Face, it is designed to teach language models to think step by step. The dataset was last updated on 2026-05-25.
The Great Barrier Reef Marine Park is the focus of this regional-scale physical dataset. It synthesizes over 3,000 surface sediment samples from Geoscience Australia's MARS database with a geomorphic features dataset, marking the first such synthesis since the 1980s. The data reveals regional trends and local-scale characteristics in sediment grain size distribution, including gravel, sand, and mud concentrations.
PhysXVerse bridges a critical gap in physics-annotated 3D datasets. It is the first general physics-grounded 3D dataset systematically annotated across five foundational dimensions: absolute scale, material, affordance, kinematics, and function description. The dataset is authored by PhysX-Omni and was last updated on HuggingFace in May 2026.
Synthetic Chinese dialogue sessions generated by the ASDAgent framework for autism intervention research. The dataset is designed for studies on ABA-aligned dialogue generation, intervention strategy modeling, and conversational agent evaluation. It was authored by 'neuljh' and last updated on May 22, 2026.
ZonMw's Zorg Vooruit+ program in the Netherlands documents the co-creation of hybrid physical-and-digital rehabilitation pathways. Over 90 PDF files from October 2021 to March 2023 capture collaborative design, development, and evaluation processes across multidisciplinary teams. The collection was authored by Marleen de Mul and harvested by DataverseNL.
Thirty cyclists participated in two 20-minute trials within a virtual reality environment. The dataset, authored by Sem Otten and last updated in May 2026, contains measurements of exerted effort, perceived effort, and psychological momentum under proximal and distal optic flow conditions.