Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
26,000 line-kilometres of total magnetic intensity (TMI) data were acquired for Geoscience Australia in 2008/2009. The processed data measures variations in the Earth's magnetic field to reveal geological structures beneath the seafloor. Quality checks were performed by GA geophysicists to ensure the data is fit-for-purpose.
A standardized and reformatted version of the original litbank-fr coreference resolution dataset provides a unified document structure. This formatting aims to simplify cross-dataset comparison, multilingual experimentation, and benchmarking of NLP systems. The dataset was uploaded by lattice-nlp and last updated on May 20, 2026.
KDV PRASAD's dataset supports the study 'From Destination Image to Digital Advocacy: A Cognition–Affect–Involvement Model of Tourists’ eWOM Intentions'. It contains empirical data from 630 domestic tourists across major destinations in India, used to develop and validate a psychological model of electronic word-of-mouth generation. The dataset was last updated on 2026-04 16 and is available under a CC-BY-4.0 license.
Annual estimates from 1990 to 2021 for all-age numbers and age-standardized rates of burden measures in the NAME region, disaggregated by sex. The dataset was created by Sina Golestani and published on figshare under a CC-BY-4.0 license. It is a 130.2 KB CSV file last updated in May 2026.
Total magnetic intensity (TMI) data measures variations in the Earth's magnetic field caused by rock-forming minerals. The processed line dataset from the GA302 survey was acquired in 2006 for Geoscience Australia and includes seismic, gravity, magnetic, and bathymetry measurements. GA geophysicists checked the data quality to ensure it is fit-for-purpose.
The Tasmania Basin hydrogeological inventory dataset contains descriptive attribute information for groundwater features in the Tasmania Basin. It covers approximately 30,000 square kilometres of onshore Tasmania and includes grouped themes such as Location, Geology, Hydrogeology, and Land Use. The dataset is provided by the Australian Ocean Data Network on data.gov.au.
321 survey responses from marathon participants in China provide data on psychological need satisfaction and event support. The dataset underpins a study using exploratory factor analysis and structural equation modeling. It validates a three-dimensional structure of need satisfaction perception.
Real property assessment records from Maryland's State Department of Assessments and Taxation and Department of Planning. The dataset includes fields for property values, tax credits, sales history, and deed references, with data updated as of March 30, 2026. Records are sourced from official state government databases.
Santa Clara County, California, provides records of deaths under the Medical Examiner-Coroner's jurisdiction from January 1, 2018 onward. The dataset includes jurisdictional and reportable non-jurisdictional cases, with nightly updates managed by the Santa Clara County government. It captures cause and manner of death determinations as mandated by California law.
Data from the Department of Agriculture documents bee assemblages in the Tuskegee National Forest, Alabama. The dataset contains 1,494 bee captures representing 38 taxa from 15 genera and 5 families, sampled from canopy and understory strata using blue vane traps from April-October 2021 and March-October 2022. It records abundance, species richness, and Simpson diversity metrics for vertical strata across two years.
Liza Hovhannisyan published a 598.5 MB dataset on figshare in April 2026. It contains all data, input files, and analysis scripts to reproduce results from a study of relativistic runaway electron avalanches. The study uses CORSIKA Monte Carlo simulations at four high-altitude stations: Aragats, Nor Amberd, Lomnický Štít, and LHAASO.
Data.ct.gov provides Connecticut's monthly online casino gaming financial data. It includes metrics like wagers, win/loss, promotional deductions, and state payments from October 2021 onward. The dataset is maintained by the State of Connecticut.
Local Nature Reserves are designated by authorities for local natural heritage and public access. The dataset for Aberdeenshire includes mandatory attributes like site name, designation date, website URL, and PA code. Data originates from NatureScot and local authorities, managed by the Government Digital Service.
An Excel template for Indonesian university lecturers to plan and document their annual performance (e-SKP). It was created by Eric Kunto Aribowo by synchronizing official rubrics from BKD, PO PAK, and e-Kinerja BKN systems. The template was last updated on April 9, 2026.
A conversational speech corpus by Liva AI (YC S25) containing over 30,000 hours of recordings from more than 8,000 speakers across 17 languages. The dataset card was last updated on 2026-05-21. The source audio was natively recorded with separate speaker channels, but samples are presented as combined conversations.
A large-scale Swahili corpus intended for continued pretraining and language exposure tasks. The dataset is maintained by NileAGI and was last updated on June 18, 2026.
Latin Classical Intertextuality Corpus contains processed texts from classical Latin authors, serving as a retrieval corpus for intertextuality research. The dataset is part of a larger collection focused on detecting intertextual relationships between Jerome (Hieronymus), Cicero and classical Latin literature. It was authored by julian-schelb and last updated on 2026-05-16.
Cenomanian (Upper Cretaceous) strata of Bathurst Island, Northern Territory, Australia, contain a newly described ophiuroid fossil. The dataset, published by Geoscience Australia Data, describes the species Ophiomicros bathursti gen. et sp.nov., which may be allied to Ophiura and Amphiura genera. The description distinguishes the new genus by its unusually large oral plates and small adoral plates.
Wide-angle seismic, gravity, and marine reflection data define crustal architecture between the Australian craton and the Timor Trough. The dataset, sourced from the Australian Ocean Data Network, was last updated in April 2026. It provides measurements of crustal thickness, P-wave velocities, and subsurface reflections along a specific transect.
A slim, flattened derived version of the Natural Questions dataset for short-answer question answering experiments. The dataset, created by BOB12311, removes original document HTML and long answer candidates to provide simple question-answer pairs. It was last updated on the platform on 2026-05-17.