Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,304 datasets
A 2017 survey collected 285 seabed sediment samples from inner Darwin Harbour and shallow water areas in and around Bynoe Harbour. The project was led by the Northern Territory Government and involved Geoscience Australia and the Australian Institute of Marine Science, funded by the INPEX-led Ichthys LNG Project. The data, including sediment samples and seagrass observations, is intended for creating baseline habitat maps to support marine resource management.
Global one-degree gridded data provides daily atmospheric profiles and surface properties derived from microwave observations. The dataset contains retrieved geophysical parameters, including temperature and water vapor profiles, surface skin temperature, and sea surface temperature, using the RAMSES II algorithm on Suomi NPP ATMS instrument data. It is an all-weather product without cloud clearing, and includes quality control flags for each variable.
Daily near-real-time model results from January 2024 and earlier for the Great Barrier Reef. This dataset contains outputs from version 3.2 of a 1km-resolution biogeochemistry and sediments model, forced by hydrodynamics and river inputs. The dataset was officially retired in February 2026 due to a model error producing unrealistically high Chlorophyll-a levels, and its authors advise discontinuing all use.
Original energy consumption data published on figshare by jiaqi Li. The dataset is available as a 10.7 KB XLSX file and was last updated on 2026-05-28. The specific source, time range, and geographic scope of the data are not detailed in the provided metadata.
Kimi-K2.6-Reasoning-3300x-WandB is a synthetic reasoning dataset containing 3,303 accepted rows. It was generated using the Kimi-K2.6 model through Weights & Biases Inference and is a subset of a larger planned 8,000-example distillation run. The dataset was created by author trjxter and last updated on 2026-05-18.
OpenS2V-Nexus is a detailed benchmark and dataset for subject-to-video generation. The released OpenS2V-5M subset contains approximately 5 million samples after filtering. It was created by author peterlrm and last updated on May 28, 2026.
ESMAtlas provides more than 700 million predicted protein structures generated by Meta AI's ESMFold language model. The dataset mirrors the ESM Metagenomic Atlas, built from protein sequences in the MGnify database. The v2023_02 release corresponds to the MGnify 2023_02 snapshot and was created using the esm.pretrained.esmfold_v1() model.
KairosQA is a question-answering dataset designed to evaluate the temporal reasoning capabilities of Large Language Models. It focuses on subjectโrelationโobject triplets from Wikidata that changed at least twice between 2018 and 2025. The dataset was created by author kyutai and was last updated on Hugging Face in May 2026.
The NSW Department of Planning and Environment acquired this multibeam survey data from the Research Vessel Bombora between 20/APR/2021 and 30/APR/2021. The dataset contains 32-bit floating point geotiff files of bathymetry and backscatter at a 5-meter resolution for the Crowdy Head Offshore area. Data processing utilized Hypack, R2Sonic GUI, POSView, POSPac, Qimera, and FMGT software.
A 2026 study by Miguel Carrasco compares human and Vision Transformer (ViT) attention patterns for 20 artisanal objects. The dataset contains results from 1,152,000 distance evaluations across four metrics, based on eye-tracking data from 30 participants and attention maps from a 12-head ViT model. It is a 5.5 KB XLS file shared under a CC-BY-4.0 license.
Areas of Interest analysis results comparing human visual attention with Vision Transformer attention mechanisms for 20 artisanal objects. The dataset contains 1,152,000 distance evaluations from four metrics, derived from eye-tracking data of 30 participants and a pre-trained DINO ViT model. Author Miguel Carrasco uploaded the 5.5 KB XLS file to figshare in April 2026.
30 participants viewed 20 artisanal objects while their gaze patterns were recorded with a Pupil Labs eye-tracker. Miguel Carrasco published this dataset in 2026, which compares these human visual attention heatmaps to attention maps extracted from a pre-trained Vision Transformer model using four metrics across 1,152,000 distance evaluations. The analysis identifies specific attention heads, such as head #12, that show the strongest alignment with human visual patterns.
A historical geography dataset reconstructing settlement evolution in Xilingol League, a pastoral region in northern China, between 1915 and 1980. The data was created by Sen Mu using declassified statistics and historical documents, employing kernel density estimation and hotspot analysis. It was last updated on April 10, 2026.
Xilingol League, a pastoral region in northern China, is the geographic scope of this dataset. It contains reconstructed settlement boundaries and patterns for the period 1915โ1980, derived from declassified statistics and historical documents. The dataset was authored by Sen Mu and last updated on 2026-04 10.
108.0 B of geospatial boundary data reconstructs settlement evolution in a pastoral region of northern China. Sen Mu's study, published on figshare in 2026, uses declassified statistics and historical documents to analyze patterns without remote sensing imagery. The research identifies a fan-shaped expansion trajectory and quantifies the influence of cultivated land and livestock as primary drivers.
A 5.1 KB ZIP file published by Zheng Zhou on figshare in April 2026. The dataset contains results from a method to calibrate Quantitative Adverse Outcome Pathways (qAOPs) on multiple chemical data, aiming to separate chemical-specific heterogeneity from core pathway effects. It includes a simulation study and a case study applying the calibration to derive points of departure for a nonmutagenic liver tumor qAOP.
A 36.6 KB Excel dataset presenting a calibration method for Quantitative Adverse Outcome Pathways (qAOPs) to separate chemical-specific heterogeneity from core pathway effects. The dataset, authored by Zheng Zhou and last updated in April 2026, includes a simulation study and a case study on nonmutagenic liver tumor qAOPs to demonstrate the approach.
An enhanced version of a binary scam classification dataset expands to 5 multi-class categories for granular detection. The original dataset contains 14,000 rows of SMS and email-style messages from an Indian context, focusing on banks, UPI, Aadhaar, and government agencies. The dataset was created by Shade63 and last updated on Hugging Face in May 2026.
Two geophysical maps of Australia at a scale of 1:25 million, showing free-air anomalies and Bouguer anomalies on land with free-air anomalies at sea. The maps were prepared from a data bank used for the 1:5 million 1976 Gravity Map of Australia, with additional marine observations from the 'Gulf Rex' vessel. They were drawn by BMR's Geophysical Drawing Office and printed by the Division of National Mapping, Department of National Resources.
Parkour-STL is a dataset of Unitree Go2 quadruped parkour rollouts generated in IsaacSim 6.0 and IsaacLab 3.0. The dataset pairs these trajectories with a library of normalized, STL-monitorable predicates sampled at 10 Hz. It was created by seil-umd and last updated in May 2026.