Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
87.3% accuracy was achieved by the PANDIA AI system for infant pain assessment across four datasets. The results, published by Oussama El Othmani in May 2026, evaluate a multimodal system combining hierarchical learning, graph reasoning, meta-learning, and symbolic explanations on data from 2,847 infants.
Jamie Davis authored a 3.6 KB header-only C++11 source code generator for deterministic dataset creation on bare-metal microcontrollers. The module, last updated on 2026-06-02, provides a zero-heap, freestanding engine using Q16.16 fixed-point precision and integrated CRC32 checksums. It is designed for hardware verification in real-time embedded systems like robotic actuators and digital signal conditioning.
NSW Department of Climate Change, Energy, the Environment and Water provides historic ecological data from environmental flows projects across New South Wales, Australia, since 1995. The data is an umbrella collection for projects like the Integrated Monitoring of Environmental Flows (IMEF) program, covering regulated rivers including the Gwydir, Namoi, Hunter, Macquarie, Lachlan, Murrumbidgee, and Barwon Darling Rivers. The dataset is stored as PDF files and was last updated on June 12, 2026.
A single electrical distribution substation in Brazil provided 1,660 images collected over two years. The dataset contains 50,705 annotated objects across 15 classes of equipment, such as disconnect switches and insulators. Images were curated by electrical engineering experts to support automated inspection tasks.
A 1992 survey by the vessel 'Rig Seismic' collected approximately 2000 km of seismic profiles and nearly twice as much bathymetric data around Christmas Island. The Australian Geological Survey Organisation (AGSO) compiled this data with other sources to produce new bathymetric and sediment thickness maps. The resulting 1:1,000,000 scale map provides more detail on seamounts and trench structures than previous compilations from the 1970s and 1980s.
3,254 Chinese reservoirs are described by 512 attributes covering upstream catchments, topography, climate, land use, and soil. The dataset also includes time series of 15 meteorological variables and satellite-derived water storage, level, area, and evaporation data for most reservoirs. Youjiang Shen from The University of Tokyo compiled this data, which is associated with a manuscript submitted to Earth System Science Data.
Cross-correlation coefficient maps for Venusian mesoscale cloud morphology, generated from images acquired by the Akatsuki orbiter's cameras at various wavelengths. The data consists of CSV files with 2880 longitude grids (0-360 degrees) and 1440 latitude grids (-90 to +90 degrees) at a 0.125-degree pixel resolution. This archive was created by researchers including Takeshi Imamura of The University of Tokyo for the paper 'Correlation of Venusian Mesoscale Cloud Morphology Between Images Acquired at Various Wavelengths' published in the Journal of Geophysical Research - Planets.
A methodology paper by Brenda López Cabrera of Humboldt-Universität zu Berlin proposes a functional data approach for forecasting generalized quantiles of electricity demand. The method is applied to load data from a transmission system operator and a balancing unit in Germany and is evaluated against other models, outperforming them in terms of mean absolute percentage error and mean squared error. Supplementary materials for the article are available online.
Present-day summaries for 77 community areas in Chicago derived from WRF weather simulations, satellite observations, and socioeconomic surveys. The dataset includes variables like land surface temperature, NDVI, median income, and Hardship Index, created by TC Chakraborty of Pacific Northwest National Laboratory. Data sources include the Chicago Data Portal and open-source WRF model simulations.
5000 ensemble members quantify the uncertainty in global mean sea-level contributions from ice sheets, glaciers, and thermal expansion. This data supplement for a 2020 Nature paper provides global, basin-mean, and regional time series alongside spatial patterns of solid-Earth deformation. Researchers can analyze the barystatic and steric drivers of sea-level change at multiple scales from 1900 onward.
Two sets of Lagrangian trajectory files generated by the TRACMASS algorithm using a 1/12-degree ocean sea-ice hindcast. The data traces the Atlantic Meridional Overturning Circulation lower limb, with trajectories initiated across the Fram Strait and the eastern Subpolar North Atlantic section. The dataset was created by Dipanjan Dey of the University of Southampton.
Exposure to ambient particulate matter is associated with up to 8.9 million deaths per year worldwide. This dataset contains minute-averaged sensor readings from a year-long field study comparing four low-cost PM sensor models deployed at two schools in Southampton, UK, from March 2018 to February 2019. It was created by Florentin M. J. Bulot of the University of Southampton and includes sensor data, meteorological wind data, and correlation analysis files.
Kilian Vos from UNSW Sydney provides 40 years of tidally-corrected shoreline change time-series for sandy coastlines around the Pacific Rim. The dataset covers approximately 3,000 beaches and over 100,000 cross-shore transects derived from Landsat 5, 7, and 8 imagery between 1984 and 2025. Data was last updated in January 2026 and was used in published research on the impact of ENSO on beach erosion and accretion.
Pacific Rim wave-dominated sandy coasts, including those in Australia, New Zealand, Japan, Chile, Peru, Mexico, and the USA, are covered. The dataset contains 40 years of tidally-corrected shoreline change data derived from Landsat 5, 7, and 8 imagery, covering over 3,000 beaches and 100,000 cross-shore transects. It was created by Kilian Vos of UNSW Sydney to investigate the impact of El Niño/Southern Oscillation on coastal erosion and accretion.
University of California, Berkeley researchers provide fluctuation electron microscopy data for comparing analysis methods on amorphous tantalum. The dataset includes raw scanning nanodiffraction data and mean CBED pattern images for both simulated and sputter-deposited 8 nm-thick Ta samples, tilted between 0 and 45 degrees. Simulated data was generated using Prismatic STEM software, while experimental data was collected on a TitanX microscope at 200 kV.
An auto-labeled dataset for fingerprinting-based indoor localization, generated via the VI-SLAM2tag method described in a 2022 IPIN conference paper. It was created by Marius Laska at RWTH Aachen University and includes raw and annotated data from an Android app, plus evaluation data and model weights. The dataset is designed to support research in low-effort labeled data collection for indoor positioning systems.
An algorithm evaluation dataset from RWTH Aachen University by Marwan Hassani, focusing on interval-based sequential patterns. The description discusses synthetic and real-world datasets used to evaluate the PIVOTMiner algorithm for mining frequent sequences and the PIVOTRanker algorithm for ranking rare patterns based on outlierness. The work addresses the challenge of finding meaningful patterns in time-interval data.
Benedikt Knopp from RWTH Aachen University created an artificial event log for container logistics. The log contains 35,761 events and 14,013 objects across 7 object types, simulated according to the OCEL 2.0 standard using CPN-Tools. It models the process from customer order registration to container shipment, including events like booking vehicles and loading trucks.
A seamless topographic color map service covering all of Australia and its external territories. The map integrates data from Geoscience Australia, the Australian Antarctic Division, and OpenStreetMap, featuring cultural, hydrography, marine, transport, vegetation, and relief themes. The topographic information was checked in 2008 and supplemented in 2009, with limited field checking noted.
Australia's continental coastline is covered by high and low tide mosaics generated from satellite imagery. The dataset likely contains pixel-based surface reflectance composites created using a multi-resolution tidal model and a Voronoi mesh to account for tidal dynamics. The data was produced by researchers including Sagar, Phillips, Bala, Roberts, and Lymburner, with a related publication in the journal Remote Sensing in 2018.