Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
OwnedByDanes's Usenet Corpus 1980–2013 is a deduplicated and sanitized archive of Usenet posts. It contains over 300 billion tokens of human discourse spanning three decades of internet history, with the full corpus comprising 103.1 billion tokens across 408 million posts. A sample of 65,000 rows is hosted on the platform.
Granger causality tests and network analysis reveal spatial correlations in standard innovation across Chinese provinces. Shuo Wang's study constructs total, positive, and negative correlation networks using data from 2001 to 2023. The analysis identifies regional disparities and key influencing factors like economic development and industrial structure.
Santa Clara County's Medical Examiner-Coroner records from January 1, 2018 onward, covering deaths under its jurisdiction and reportable non-jurisdictional cases. The dataset includes columns for cause and manner of death, demographics, and incident location. It is updated nightly and hosted by data.sccgov.org.
litbank-en is a reformatted version of the original coreference resolution dataset for literature. The dataset is standardized to provide a unified document structure for cross-dataset comparison and multilingual experimentation. It was created by lattice-nlp and last updated on 2026-05-20.
Olena Zhabenko's dataset contains results from a network analysis study of 661 adult survivors of traumatic events. It examines relationships between rumination, measured by the RTQ-10, and PTSD symptoms, measured by the PCL-5, while controlling for depression (PHQ-9) and anxiety (GAD-7). The analysis reveals selective associations at both cluster and item levels.
661 adult survivors of traumatic events participated in an online study, completing standardized questionnaires for PTSD, rumination, anxiety, and depression. The dataset supports a multi-level network analysis examining granular relationships between specific PTSD symptoms and rumination while controlling for comorbidities. It includes responses from participants with a median age of 35 years, 49.3% of whom were female.
A paper summarizing the stratigraphy, distribution, and geometry of mid-Tertiary units forming subsurface permeability barriers in the Murray Basin. The analysis is based on subsurface facies analysis of borelogs and palaeogeographic reconstructions, indicating at least three separate marine incursions during the Cainozoic. The data originates from Geoscience Australia and was last updated in March 2026.
Quebec's public administration publishes monthly updated data on the health status of its information resource projects. The dashboard, first published in December 2020, is maintained by the Government and Municipalities of Québec.
Ricoh AI developed the JDocQA Reasoning Bench to evaluate the reasoning performance of multimodal large language models on documents containing figures and charts. The benchmark was created as part of Japan's GENIAC project, which aims to strengthen domestic generative AI development capabilities. It was published on the Hugging Face platform by the author ricoh-ai.
EricLu created a dataset of scientific problems paired with solutions. A newer version, SCP-378K, contains 377,705 examples and provides extracted source solutions for every problem. The dataset was last updated on May 8, 2026.
Geoscience Australia Data provides a detailed scientific description of the Great Cumbung Swamp, the unique terminus of Australia's low-gradient Lachlan River. The report describes three distinct depositional environments, including a sinuous channel up to 40 meters wide, a vast Phragmites marsh, and overflow areas with channels up to 20 meters wide. It was last updated on March 25, 2026.
ArithMark 2.0 is a procedurally generated benchmark for evaluating integer arithmetic ability in language models. The benchmark is designed for base-model log-likelihood scoring and does not require instruction following or chain-of-thought reasoning. It was created by AxiomicLabs and was last updated on the Hugging Face platform in May 2026.
Hourly performance readings for Combined Heat and Power, Fuel Cell, and Solar Electric systems are collected from New York State sites. The New York State Energy Research and Development Authority (NYSERDA) maintains this database, which is updated nightly with data from the previous day. It includes information on systems funded by NYSERDA and all operational energy storage systems in the state.
NYSERDA's Distributed Energy Resources Integrated Data System contains performance data for technologies like Combined Heat and Power, fuel cells, and solar electric systems larger than 50 kW. The system provides hourly readings for each site, with historical databases updated nightly to include the previous day's data. The New York State Energy Research and Development Authority maintains this web-based system, which includes data beginning in 2001.
A 1034 sq km backscatter grid of the Lord Howe Island shelf, produced from processed EM300 sonar data collected during a 2008 marine survey. Geoscience Australia conducted the survey using the RV Southern Surveyor, merging bathymetric data with pre-existing datasets. The survey characterized benthic environments through sediment sampling, video observation, and oceanographic measurements.
A 1034 sq km backscatter grid details the seabed composition around Lord Howe Island. Geoscience Australia collected the data during a 2008 marine survey using the RV Southern Surveyor, merging it with pre-existing bathymetric data. The survey also involved sediment sampling, rock coring, underwater video, and oceanographic measurements.
An API provides developer access to the underlying transport planning data used by Translink's Journey Planner, website station screens, and mobile application in JSON format. The data is provided by the Government Digital Service under an Open Government Licence, which requires acceptance of an Opendata API Licence and user agreement.
Geoscience Australia's South and Southwest Regional Project conducted a regional seafloor mapping study during 2000/2001. The work delineated four major geomorphological features and defined five acoustic echo facies for the Great Australian Bight area. The results were digitized into a GIS to support biological, environmental, and economic assessments for regional marine planning.
Five distinct acoustic facies, representing seabed sediment types from undisturbed layers to disturbed deposits, were mapped across the Great Australian Bight margin and abyssal plain. The dataset captures geomorphological features including continental shelves, slopes, terraces, and canyons within a Geographical Information System. This regional seafloor mapping study was conducted by Geoscience Australia during 2000/2001 to support marine planning.
Geospatial data describes the Harris Greenstone Belt, a late Archean-Proterozoic terrane in the Gawler Craton. It includes interpreted distributions of rock types like komatiite, basalt, and banded iron formation. The dataset was published by the Australian Ocean Data Network and was last updated in April 2026.