Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,416 datasets
Experimental results for various textured tools, measuring Specific energy consumption, Specific cutting pressure, Tool-tip temperature, and Material Removal Rate (MRR). The dataset was authored by Palanisamy Angappan and last updated on 2026-05-12. It is a small dataset of 9.5 KB, stored in an XLS file.
2026 data from Canada's official greenhouse gas inventory, submitted to the UNFCCC. The dataset is provided by Environment and Climate Change Canada and includes the National Inventory Report and supporting data files. The material is available in various formats, including HTML documents and downloadable data files.
2,535 audio utterances in Yucatec Maya (ISO 639-3: yua) with orthographic transcriptions, compiled for training and evaluating automatic speech recognition models for low-resource indigenous languages. The corpus was created by author mau-cr as part of a thesis on MMS-based ASR for Yucatec Maya and was last updated on Hugging Face in May 2026.
A dataset of 49,660 parameterized 3-D computational fluid dynamics (CFD) surface meshes generated by IBM Research. Each sample is a distinct geometry from a 17-parameter family and includes per-vertex pressure and wall shear stress fields. All samples share the same triangulation topology of 9,600 triangles and 25,600 vertices.
IceCube's 10-year sample of track-like neutrino candidate events used for point-source searches, covering data from April 2008 to spring 2018. The dataset includes reconstructed particle information, detector uptime lists, and instrument response functions for each season. It was created and released by the IceCube Collaboration.
An essay by Konrad Hinsen of the Centre National de la Recherche Scientifique, published via Open Access. It discusses the concepts of replicability, reproducibility, and verifiability in scientific research, with a focus on computer-aided work. The text argues for verifiability as a crucial link between replicating observations and reproducing conclusions.
Supplement Material is a document authored by Ren-Hui Zheng, published on figshare under a CC-BY-4.0 license. The 23.5 KB file, last updated on 2026-05-22, likely contains supplementary information for a rate equation model describing energy transfer processes. The specific data format and content require verification after download.
INTEGRAL Observing Program data contains pointed observation targets from the ESA INTEGRAL gamma-ray astronomy mission. The HEASARC database table includes targets from both the Core Program (Guaranteed Time) and the General Program (Open Time) accepted observations, covering Announcement of Opportunity cycles AO-1 through AO-20. NASA's HEASARC maintains the table, which was last revised in August 2007, updated to include AO-20 in November 2022, and is refreshed weekly from the ESA mission website.
Norfolk's complaint tracking system contains records of citizen-reported and staff-reported issues handled by the Department of Neighborhood Services and City Planning. It includes a rolling three years of data updated daily, covering property issues like tall weeds, trash, abandoned vehicles, and graffiti. The data is provided by the City of Norfolk via data.norfolk.gov.
Protein interaction data for alpha-synuclein monomers and fibrils, generated via micromap photoproximity labeling in mouse brain lysate. Marshall G. Lougee published the dataset in April 2026. The data includes validation against prior proximity labeling sets and further investigation examples.
Between May 29 and June 19, 2017, benthic sediment sampling was conducted in inner Darwin Harbour and shallow water areas around Bynoe Harbour. The dataset comprises grain size measurements on seabed sediments, collected as part of a four-year (2014-2018) science program to establish baseline data for thematic habitat maps. This work was led by the Northern Territory Government in collaboration with Geoscience Australia, the Australian Institute of Marine Science, and supported by the INPEX-led Ichthys LNG Project.
RoMo-SMPLX is a large-scale dataset of single-person body motion sequences in the SMPL-X parameter space, paired with rich multi-level text descriptions. It is the raw representation of the RoMo body corpus, with the same motions also released in other feature formats. The dataset is authored by RoMoDataset and was last updated on May 21, 2026.
Geospatial boundaries for protected bathing water areas under the EU Bathing Waters Directive. The dataset was generated by the Government Digital Service using Land and Property Services (LPS) land parcel line boundaries and a 1:10k raster basemap. Beach boundaries are defined by land parcel lines at the upper extent of the beach and the Low Water Mark of Mean Tide (LWMMT) line.
Deepfabric is a tool for generating, training, measuring, and evaluating high-quality synthetic data within a single pipeline. The project is authored by nolabs-ai and was last updated on June 23, 2026. It is distributed under the Apache-2.0 license.
Bathing Waters Directive - Protected Areas is a geospatial dataset from the Government Digital Service. The shapes are generated using Land and Property Services land parcel line boundaries and a 1:10k raster basemap. Beach boundaries are defined by land parcel lines at the upper extent of the beach and the Low Water Mark of Mean Tide line.
Bathymetry data from a contracted national reference survey in Hobart, Tasmania, acquired for the Australian Hydrographic Office between 10 October and 8 November 2021. The surface was created for calibrating multibeam echosounders and is provided as separate 0.5m and 1m resolution grids in MSL, LAT, and Ellipsoid vertical datums. The dataset was processed using QPS Qimera and exported as 32-bit floating point GeoTIFF grids.
234,028 individuals from the UK Biobank underpin the SpiroLLM model for generating diagnostic reports from pulmonary function tests. The model extracts morphological features from respiratory curves and aligns them with PFT numerical values. It achieved a diagnostic AUROC of 0.8977 and maintained a 100% valid response rate in robustness tests.
A September 2020 bathymetry survey acquired for the Australian Hydrographic Office in Gulf St Vincent, South Australia. The dataset consists of 1m resolution grids for two surveyed sites, processed from multibeam and singlebeam echosounder data. It was created for calibrating multibeam echosounders under the Hydroscheme Industry Partnership Program.
NASA's ACE spacecraft provides an ensemble of 26 data sets of hourly-averaged, background-corrected energetic particle fluxes. The data includes fluxes for protons, electrons, and ions with Z>1 across multiple energy bins, measured in different telescope apertures and reference frames. The data was last updated on 2026-03 -13.
NASA's ACE EPAM dataset is an ensemble of 26 data sets providing 17-minute-averaged, background-corrected energetic particle fluxes. The data includes fluxes for protons, electrons, and ions across multiple energy bins, measured in different frames and from various telescope apertures. The data was last updated on 2026-03 13.