Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
A bilingual English and Japanese judgment-eliciting Q&A corpus encodes the documented judgment from four research lines in the shimo4228 research program. The corpus is released under CC0 to maximize LLM-mediated diffusion and serves as an operational form of a specific authorship strategy tactic. It was authored by Shimo4228 and last updated on May 22, 2026.
Computational chemistry data examines the potential energy surface for a covalently bound N2H6 dimer. The dataset contains results from model chemistry calculations probing the structure and stability of this high-energy local minimum. Author Kelling J. Donald published this work on figshare in April 2026.
One supplementary document presents a theoretical argument that causation drives changes in thermodynamic entropy, time progression, and measurable information evolution. The 695.5 KB DOCX file was authored by Georg Franz Weber and uploaded to figshare in April 2026. It synthesizes concepts from physics and information theory to propose a foundation for defining knowledge.
National Gravity Compilation 2019 Ausdrape elevation geoid image (hillshade HSI) is a digital elevation dataset for Australia. The data were processed from SRTM surface elevation 3 second grid data and vertically continued to a drape surface with a minimum clearance of 250 meters. The grid has a cell size of approximately 435 meters, and the data are provided in units of meters.
Orbits 6-350 of magnetic field data from the CRRES satellite's tri-axial fluxgate magnetometer, converted to sensor coordinates. The dataset was processed by NASA using software and calibration coefficients provided by the principal investigator, Dr. Howard Singer. Original raw files for these orbits were not readily available, so the converted files are provided for uniformity, along with source code and calibration files.
Flood extent polygons derived from satellite imagery for emergency response in selected Canadian regions. Natural Resources Canada maintains this archive of all flood products generated since 2005, updated in near real-time during events. The dataset includes products validated on a best-effort basis.
Labeled cohort splits used to evaluate CGM encoders on two binary metabolic outcomes β insulin resistance and Ξ²-cell dysfunction. The dataset was created by CRUISEResearchGroup for the paper CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretraining. It was last updated on HuggingFace on 2026-05-13.
Aerospace Corporation compiled raw telemetry data from eight scientific instruments on the CRRES satellite, including particle spectrometers and a photometer. The data covers specific orbits from the mission, with spacecraft ephemeris, attitude, and magnetic field information provided in separate files. NASA is the organization responsible for this dataset, which was last updated in March 2026.
A 2022-2023 seabed survey acquired by the New South Wales Department of Planning and Environment onboard the RV Bombora. The dataset contains 32-bit floating point GeoTIFF files of bathymetry and backscatter at a 5-meter resolution for the Solitary Islands Marine Park area. Data was processed using Hypack, R2Sonic GUI, POSView, POSPac, Qimera, and FMGT software.
FDAbench-Full is a benchmark dataset designed for evaluating data agents in multi-source analytical scenarios. It contains 2,007 diverse tasks across different data sources, domains, difficulty levels, and task types. The dataset was created by FDAbench2026 and was last updated on May 20, 2026.
Reconnaissance-style maps of surficial cover facies on the Great Barrier Reef, compiled in 1982 by the Bureau of Mineral Resources. The maps apply a simple bathymetric classification differentiating supratidal, intertidal, and subtidal zones. Cartography was completed by G.A. Young and R.A. Swoboda, with editing in 1983.
General statistics for the draft genome assembly and in silico annotated gene models of the organism E. neotenicus. The dataset is a 10.5 KB Excel file authored by Natan HorΓ‘Δek and last updated in April 2026. It is shared under a CC-BY-4.0 license on the figshare platform.
Bathymetry and backscatter data for the North Wollongong coastline from Bellambi Point to Stanwell Park, NSW. The dataset was acquired by the New South Wales Department of Planning and Environment using an R2Sonic 2022 multibeam sonar onboard the RV Bombora between August 2017 and March 2022. It consists of 32-bit floating point GeoTIFF files at 5-meter resolution, processed from raw sonar data.
NSW Department of Planning and Environment acquired bathymetry data from Bellambi Point to Stanwell Park. The survey was conducted from August 2017 to March 2022 using an R2Sonic 2022 multibeam sonar aboard the RV Bombora. This dataset provides 5-meter resolution geotiff files of seabed depth and backscatter intensity.
A 5.5 KB Excel file compares the performance of different language models as the basis for a joint training method. Authored by Pegah Safari, the dataset was last updated on May 11, 2026. The specific models, metrics, and number of rows are not detailed in the available metadata.
Thai Lexical Decision Task (42 items) score distributions from online and interview samples. The dataset, authored by Graham Pluck, was last updated on May 4, 2026. It is a 5.5 KB Excel file containing psychometric properties and correlates.
Five years of city contract data, updated monthly, lists open and closed agreements. It includes agency, supplier, dollar amount, procurement type, and contract start and end dates. The dataset is provided by data.richmondgov.com.
3,378 multilingual conversations in OpenAI chat format, last updated on 2026-05-21. The dataset contains ADHD-support dialogues between a user persona and an AI companion modeled on Dr. Ned Hallowell's teachings. It was created by Reza2kn and covers 69 ADHD-related issues across nine languages.
3,347 child deaths in England from 1 April 2019 to 31 March 2020, equating to approximately 28 deaths per 100,000 children. The report was commissioned by the Healthcare Quality Improvement Partnership (HQIP) on behalf of NHS England and is published by the Government Digital Service. It analyzes characteristics of these deaths to inform future improvements in child health outcomes.
Voyager 1's Low Energy Charged Particle experiment recorded intensities of electrons and ions during its close encounter with Jupiter. The dataset includes nearly 100 instrument channels, providing 48-second rate and flux measurements for particles like protons, alpha particles, and nuclei. NASA produced this calibrated data, which was last updated on March 13, 2026.