Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,463 datasets
The O. D. Skelton Memorial Lecture series honors Oscar Douglas Skelton, a key architect of early Canadian foreign policy. Inaugurated in 1991, it features distinguished speakers examining topics related to Canada's foreign policy, international development, and trade. The series is coordinated by Global Affairs Canada's Open Insights Hub.
The Commission on Isotopic Abundances and Atomic Weights (CIAAW) of IUPAC completed its last update of isotopic compositions in 2009. This table presents evaluated data from the 'best measurement' of isotope abundances for each element, along with representative abundances and uncertainties for normal terrestrial materials. The data is consistent with the standard atomic weights recommended by CIAAW in 2007.
UK Environment Agency data defines the hierarchical relationships between Water Framework Directive (WFD) classification items. The spreadsheet structures elements, components, and sub-elements used to calculate headline waterbody quality scores. Data is sourced from the Environment Agency's Catchment Planning System and was last updated in July 2026.
Three synthetic data sets generated for cross-verifying parameter estimation codes used in NICER X-ray astronomy analyses. The data simulates observations of millisecond pulsar PSR J0030+0451, including scenarios with ultra-compact stars and single or dual hot spots. The deposit includes auxiliary files and checksums, created by Slavko Bogdanov of Columbia University.
The Annotated Corpus of Classical Tibetan (ACTib), Part I is a part-of-speech tagged version of the Buddhist Digital Resource Center's digitized Tibetan etext collection. It was created using a memory-based tagger trained on a separate POS-tagged corpus of Classical Tibetan. The corpus includes files that were not manually corrected and contains some annotations from corrupted source files.
Simulated datasets from the Large Hadron Collider used to hunt for parity-violating signals beyond the Standard Model. The data, created by C. G. Lester at the University of Cambridge, contains reconstructed four-momenta for the five hardest jets per event, stored in HDF5 files with shape (n, 20). Truth-jet files include additional truth-level flavour and helicity information encoded using PDG IDs.
A simulated dataset of backscattered waves governed by the 2D wave equation from randomly placed Dirichlet particles. The data was created by Artur L. Gower of the University of Manchester and includes subsets for incident waves with different wavenumber (k) and time (t) ranges. The dataset is provided via the paperswithcode platform.
Salford City Council Senior Salaries provides detailed information on the remuneration and roles of senior council staff earning over Β£50,000 annually. The dataset includes salary brackets, job titles, responsibilities, and details on bonuses and benefits-in-kind, compiled to meet the UK Local Government Transparency Code 2014. It specifically identifies by name any employee earning Β£150,000 or more.
ManySStuBs4J contains single-statement bug fixes mined from open-source Java projects on GitHub, classified into 16 syntactic templates called SStuBs. The dataset has two variants: one mined from 100 Java Maven Projects and another from the top 1000 Java Projects. Bug commits were identified using the SZZ heuristic and keyword filtering in commit messages, achieving an estimated 94% accuracy.
Roundtrip is a collection of benchmark datasets compiled for evaluating deep generative neural density estimators. It aggregates multiple established datasets from the UCI repository and other sources, covering domains such as activity recognition, protein structure, finance, audio, and particle physics. The collection also includes standard computer vision datasets like MNIST and CIFAR-10, as well as datasets for outlier detection tasks.
Supplementary data for a 2016 paper published in the Journal of Physical Chemistry C. The files contain experimental data and verified computational models for hydrogen oxidation and evolution reactions. The data was authored by Anthony Kucernak of Imperial College London.
Apprenticeship vacancy and application data from the official online system managed by the National Apprenticeship Service in England. The dataset tracks vacancies posted and applications made through the Apprenticeship vacancy online system, which allows searches by geography, occupation, job role, and keywords. It was last updated on 2026-07-08 by the Skills Funding Agency under the OGL-UK-3.0 license.
FE data library: apprenticeship vacancies provides data on the number of vacancies posted and applications made through the official Apprenticeship vacancy online system managed by the National Apprenticeship Service. The data covers the period from 2008/09 to 2014/15 and is used to monitor the performance of the online portal. It specifically tracks online applications via Find An Apprentice and does not represent total apprenticeship starts or offline applications.
A 2018 baseline study for Northern Ireland provides a high-level preliminary vulnerability assessment of coastal erosion risk. The assessment consists of two stages: an Erosion Risk Appraisal layer and a separate Vulnerability Assessment comparing erosion risk against asset values. This dataset represents the foundational risk appraisal layer from that study, identifying areas potentially vulnerable to coastal erosion.
A geospatial dataset provides a high-level vulnerability assessment of physical assets along the Northern Ireland coast. It was created as part of a 2018 baseline study and gap analysis for coastal erosion risk management, prepared by Amey Consulting with HR Wallingford for government departments. The assessment compares areas of high, medium, and low erosion risk against the value of physical, historic, and natural assets.
A GIS analysis using 6 biophysical variables classified the ocean floor into 53,713 polygons across 11 seascape categories. The study, associated with the Australian Ocean Data Network, aimed to identify potential high seas marine protected areas. Validation was performed by comparing the seascapes with an existing seafloor geomorphology map.
Australian coastal regions are covered by 10m neutral stability wind speed and direction data derived from Sentinel-1 A, B, and C Synthetic Aperture Radar (SAR) satellites. The dataset uses a consistent CMOD5N geophysical model function and variational Bayesian inversion, with wind speeds calibrated against a Metop scatterometer database. Data is presented in delayed mode and derived from the Sentinel-1 Ocean wind level-2 product.
91 days of AIS vessel tracking data from Tokyo Bay between July 29 and October 27, 2024, published by Moritz HΓΌtten in 2026. The dataset includes raw message streams and processed geospatial rasters of vessel density, speed, bearing, and inferred berth areas. It supports high-resolution analysis of maritime traffic patterns at a 1-arcsecond spatial resolution.
Provincial Trails in Alberta are established under the Public Lands Act through Ministerial Order. This dataset from the Government of Alberta includes trail name, legislation, and designation information for mapping and reference purposes. The data was last updated on 2026-06-29.
A City of Melbourne mapping dataset identifies laneways with potential for four greening responses: forest lanes, farm lanes, park lanes, and vertical gardens. It uses a gridcode system where lower values indicate higher greening potential. The dataset is part of a larger municipal initiative to combat urban density, heat island effects, and biodiversity loss.