Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,171 datasets
Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual corpus for long chain-of-thought reasoning, created by lightonai and released in 2026. It spans six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, with sequences up to 32,768 tokens.
11,991 high-quality computer science literature resources and coding tasks curated by pheonix-delta. The dataset is designed to train code generation models and software agents, focusing on algorithmic logic and computer science principles. It was last updated on Hugging Face on May 24, 2026.
219,647 mathematical instruction-response pairs curated, deduplicated, and unified by author pheonix-delta. The dataset is designed to train language models on step-by-step problem-solving and was last updated on 2026-05-24.
A psychometric dataset comparing large language models to humans on the Moral Identity Scale. The data includes scores for 57 LLM variants from 9 model families under two prompting conditions, compared against a human sample of 329 individuals. Davin Nabizadehchianeh published the dataset via Harvard Dataverse in May 2026.
DOB NOW: Build โ Limited Alteration Applications dataset contains records for minor construction work in New York City, such as plumbing repairs and oil burner installations. The data includes columns for Filing Status Name, Job Number, Proposed Work Summary, Location BIN, and geographic coordinates. It is published by data.cityofnewyork.us and was last updated in April 2026.
2,738 student-instructor question-answer pairs for Calculus 1, augmented with relevant lecture notes and exercises as context. All answers were written by an experienced calculus teacher and the dataset was originally in French before translation to English using GPT-4o. The dataset was created by Jeremmmyyyyy and last updated on 2026-05-27.
PlanningBench is a synthetic planning benchmark and data construction framework for evaluating and training large language models. The dataset, created by Tencent, focuses on complex, text-based planning tasks that require coordinating goals, constraints, and resources. It was last updated on May 28, 2026.
An English language corpus sample spanning the years 1800 to 1875, hosted on Hugging Face. The dataset was uploaded by author haykgrigorian and last updated on July 6, 2026. Its specific size, format, and internal structure are not detailed in the available metadata.
A sample pack of enterprise-grade, high-fidelity human voice recordings. The dataset is optimized for low-latency conversational AI interfaces and intent-mapped software pipelines. It was authored by MarieDeVox and last updated on 2026-05-26.
Replication files for the paper "Signal Generation Neglect: Experimental Evidence on Belief Updating from Threshold Signals" by Kamijo, Taguchi, and Tsuruta. The data is hosted on Harvard Dataverse and was last updated on June 27, 2026. It likely contains experimental data on how individuals update beliefs based on threshold-based signals.
324 prenatal MRI scans from normal singleton pregnancies were used to quantify sylvian fissure development between 22 and 38+6 weeks of gestation. The study, published on figshare, established reference ranges for 10 morphological parameters like insula depth and sylvian fissure width. These normative charts successfully identified subtle cortical malformations in two genetically confirmed cases of polymicrogyria.
Talker-T2AV-Data is a clean training data package for the Talker-T2AV model, which performs joint talking audio-video generation using autoregressive diffusion modeling. The dataset is hosted by HKUSTAudio and was last updated on 2026-05-24. It contains metadata and grouped archive shards for audio, motion, and video modalities.
A guide by Yiny Paola Cรกrdenas Rodrรญguez provides concrete, operational procedures for using generative AI to energize specific moments in learning. The 12.8 KB Excel file was last updated on May 7, 2026. It is shared under a CC-BY-4.0 license on figshare.
Geospatial data details the locations of proposed and installed cable interconnections for offshore wind energy facilities. The dataset is maintained by the U.S. Department of the Interior and was last updated in April 2026. It is provided by project lessees for planning review and permitting purposes.
Offshore Wind - Cable Interconnections data from the U.S. Department of the Interior shows locations of proposed and installed facilities where wind farm power will be injected. The dataset includes multiple proposed configuration options within project design envelopes and is updated as new public information becomes available. Data was last updated in April 2026.
A database supporting an academic article on race-gender intersectionality in Mexican digital news coverage of Kamala Harris. The dataset was created by Edrei รlvarez-Monsivรกis and is hosted on figshare. It is available as a 77.9 KB XLSX file under a CC-BY-4.0 license.
12.5 KB of supplementary materials from a study on prompting with empathy. The artifacts include prompts, comparison templates, and questionnaires, shared under a CC-BY-4.0 license on figshare. The materials were last updated in May 2026.
226 patients from the Swiss IBD cohort study were examined, including 78 with documented venous thromboembolism (VTE) and 148 age- and sex-matched controls. The dataset contains quantified serum levels for 12 proinflammatory cytokines and detailed disease characteristics. The study, published on figshare under CC-BY-4.0, found a significant association between elevated IL-22 levels and VTE in IBD patients.
1.7 KB of statistical results from linear and segmented regression analyses. The table provides raw p-values, FDR-adjusted q-values, and log-likelihood ratio test results for nonlinearity, authored by Shuang Deng and last updated on 2026-04-29. It indicates significant and suggestive nonlinearity based on q < 0.05 and p < 0.05 thresholds.
Seabird concentration data highlights the most vulnerable aggregations for each month of the year in offshore waters based on abundance and vulnerability to oil pollution. The data was compiled in 1995 from the European Seabirds at Sea Database, with points representing the center of 1/4 ICES Rectangles. It is provided by the Joint Nature Conservation Committee and visualized on Natural England's MAGIC site.