Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
Statistics Canada provides trade in goods data categorized by exporter characteristics. The dataset likely contains export values and counts of exporting enterprises, segmented by enterprise employment size and the number of partner countries. It was last updated on April 24, 2026.
The Department of Water and Environmental Regulation provides geospatial data on the extent of major flood events in Western Australia. The dataset includes polyline features for specific flood events dating back to 1830, with detailed records for 11 listed major events from 1979 to 2023. It is part of a three-layer floodplain mapping series requiring sequential loading.
New York Power Authority (NYPA) data provides megawatt-hours produced net of station service by each of its 16 generating facilities. The dataset begins in 2013 and is published by data.ny.gov. It was last updated on April 14, 2026.
The City of Tempe, Arizona, maintains this compilation of address point data for all occupiable units and official addresses within its jurisdiction. The dataset includes a point location and the official address defined by the Building Safety Division, with additional attributes that may be populated. It is published weekly as the system of record, with the last update on 2026-03 21.
MahaveerAI created this multilingual dataset to train and evaluate its Bol-AI model for understanding the history of Chhatrapati Shivaji Maharaj and the Maratha Empire. The dataset is designed for AI processing and question-answering across multiple languages. Its last recorded update was on 2026-05-18.
2.4 MB of source data for the manuscript 'Theoretical quantitative model and clinical outcome predictions of conductive cardiac patches for electrophysiological treatments'. The mass spectrometry proteomics data are deposited in the ProteomeXchange Consortium via the PRIDE repository with identifier PXD073006. The dataset was authored by Yuchen Miao and last updated on April 23, 2026.
Geoscience Australia and Queensland Fire Department developed a Probabilistic Tsunami Hazard Assessment for the Gladstone region from Agnes Waters to Yeppoon. The report details modeling validated against three historic tsunami events and provides conservative inundation zone estimates corresponding to Australian tsunami warning categories. These results were published in a PDF report on April 16, 2026.
766,987 records form a large-scale instruction-tuning dataset for code generation, assembled from multiple open sources and deduplicated. The dataset was created by Voidreaper2026 and is formatted as JSONL in the ShareGPT conversational style. Its last recorded update was on May 6, 2026.
Geoscience Australia conducted a regional mapping program addressing stratigraphic and structural exploration risk. The data pack comprises seismic horizon grids, isochron grids, and fault maps for the Triassic succession of the Roebuck Basin. Mapped horizons were tied to wells using synthetic seismograms and placed within a regional tectonostratigraphic framework.
A sedimentological analysis of seabed samples, shear-stress modelling, and three-dimensional acoustic imaging reveals Keppel Bay's sediment transport pathways. The dataset likely contains geospatial data on sediment starvation zones, tidal sand ridges, and subaqueous dunes in a macrotidal embayment linking the Fitzroy River to the Great Barrier Reef shelf. It is hosted by the Australian Ocean Data Network and was last updated on 2026-04-16.
The OxyCom dataset contains psychological and hormonal data from an experimental study comparing online and offline communication in dyads of strangers. Data was collected from 131 participants in Saint-Petersburg and Krasnoyarsk. The dataset was authored by the SCILa lab and last updated in April 2026.
Congbobanan Toaan Gov Vn is a document-level mirror of the Vietnamese Cα»ng cΓ΄ng bαΊ£n Γ‘n portal from the Supreme People's Court of Vietnam. Each case is provided as a raw PDF, parsed markdown, a structured JSON record with metadata and NER, a 2,048-dimension dense embedding, and a 2-D projection with cluster IDs. The corpus was produced end-to-end by author tmquan and was last updated on HuggingFace in May 2026.
Lithostratigraphy, grain sizes, and down-hole logs from two Ocean Drilling Program sites reconstruct glacial processes in eastern Prydz Bay. The record indicates repeated advances and retreats of the Lambert Glacier-Amery Ice Shelf system from the Pliocene to Pleistocene. Data from the Australian Ocean Data Network shows the grounding line did not extend to the shelf break after 0.78 Ma.
1,102,568 high-scoring StackOverflow question-and-answer pairs used to create pre-built FAISS and BM25 indexes. The dataset, created by author 'korunil', is optimized for programming-related retrieval tasks and was last updated on 2026-05-18.
500 stratified documents form a public sample drawn from the French Premium Web Corpus v1.4.0. FINALEADS LLC built this compliance-ready training dataset from 2.66 billion tokens of French finance, regulatory, and economic open data. The dataset was last updated on 2026-05 -19.
550,000 square kilometers of descriptive hydrogeological data for the Carpentaria Basin in north-eastern Australia. The dataset groups attributes into themes like Location, Geology, Hydrogeology, and Groundwater management. It is provided by the Australian Ocean Data Network via data.gov.au.
The Money Shoal Basin dataset provides descriptive attribute information for groundwater features in northern Australia, primarily offshore in the Arafura Sea. It is grouped into themes including location, demographics, physical geography, geology, hydrogeology, groundwater management, land use, and scientific stimulus. The dataset was published by the Australian Ocean Data Network and last updated on April 16, 2026.
From 01 November 2024 to 30 June 2025, this dataset discloses contracts valued over $10,000 for the Queensland Department of Housing and Public Works. It was published by the department under the CC-BY-4.0 license and last updated in April 2026. The data is provided as part of the Queensland Government's contract disclosure guidelines.
26,000 line-kilometres of total magnetic intensity (TMI) data were acquired for Geoscience Australia in 2008/2009. The processed data measures variations in the Earth's magnetic field to reveal geological structures beneath the seafloor. Quality checks were performed by GA geophysicists to ensure the data is fit-for-purpose.
A standardized and reformatted version of the original litbank-fr coreference resolution dataset provides a unified document structure. This formatting aims to simplify cross-dataset comparison, multilingual experimentation, and benchmarking of NLP systems. The dataset was uploaded by lattice-nlp and last updated on May 20, 2026.