Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,416 datasets
Historical passenger and freight transport data for 38 and 43 countries respectively, from 1990 to 2018, sourced from the International Transport Forum. The dataset integrates World Bank socioeconomic indicators and projects future demand using the Shared Socioeconomic Pathways (SSPs) scenario framework. It was compiled by a researcher from Carnegie Mellon University for scenario analysis.
The Soundscape Attributes Translation Project (SATP) dataset provides 27 standardized 30-second binaural audio recordings from urban public spaces in London, plus a calibration signal. It includes perceptual responses from thousands of participants across multiple languages, collected via headphone-based listening experiments at partner institutions worldwide. The dataset was created to validate translations of soundscape attributes and offers a standardized set for cross-cultural perceptual research.
130GB of compressed data lists anomaly-free charge assignments for the chiral fermionic content of the Minimally Supersymmetric Standard Model plus three right-handed neutrinos. The dataset, produced by B. C. Allanach of the University of Cambridge, includes solutions up to a maximum charge magnitude of Qmax=10. It is accompanied by C++ source code for generating the solutions, Mathematica notebooks for analytic parametrization, and filter programs for subset selection.
GRIT is a global river network representing tributary and distributary components, including multi-thread rivers, canals, and delta distributaries. It is the first global hydrography produced at 30m raster resolution, created by merging Landsat-based river masks with elevation-generated streams and using the FABDEM digital terrain model. The dataset, authored by Michel Wortmann at the University of Oxford, provides vector data with network topology in GeoPackage format.
A corpus developed for the IEEE-AASP ACE Challenge to evaluate algorithms for blind acoustic parameter estimation from speech. It includes anechoic speech recordings from TU Delft and room impulse responses (RIRs) and noise recordings from seven rooms at Imperial College London. The dataset is described in a 2016 journal paper by Eaton et al. and includes multi-channel microphone configurations.
430 rows of ant community data collected as part of the SAFE research project. The dataset includes a site x species matrix with abundance counts for genera like Diacamma and Pheidole across two forest types. It was authored by Ross Gray of Imperial College London.
927 rows of morphological measurements and 1082 rows of abundance data for 339 morphospecies of ants, collected as part of the SAFE research project in Sabah. The dataset was authored by Tom R. Bishop from Imperial College London. It includes detailed trait measurements and abundance counts from soil pits across a forest disturbance gradient.
July 2016 data on Community Safety Partnerships in England, listing their names and associated codes. The dataset is provided by the Office for National Statistics under the OGL-UK-3.0 license. It was last updated on the platform on 2026-07 08 13:32:16.588935.
Antarctic GPS vertical deformation time series from 1979 to 2022, derived from seven surface mass balance (SMB) model products. The dataset was created by the Australian Ocean Data Network to evaluate model performance, computing elastic displacements using the Regional ElAstic Rebound calculator (REAR, v1.5). Data were processed into a common 2 km resolution grid and span from 1980 to 2022.
Antarctic fast ice at Cape Evans was the site for this 3-terabyte dataset collected during two field campaigns in November/December 2018-2019. It consists of high-resolution under-ice imagery from a custom sled system, ice core scans, and auxiliary data like irradiance measurements and water samples for chlorophyll-a analysis. The data were gathered by an AGP and NZARI collaboration to test an under-ice surveillance system for mapping sea-ice microbial habitats.
Stereo-Baited Remote Underwater Video Stations (BRUVS) were deployed in the proposed Oceanic Shoals Commonwealth Marine Reserve in the Timor Sea. A total of 56 BRUVS were deployed between 31 and 77 meters depth for one hour each, yielding 56 hours of video footage. The survey was conducted between 12 September and 6 October 2012 by a research collaboration including the Australian Institute of Marine Science, Geoscience Australia, the University of Western Australia, and the Museum & Art Gallery of the Northern Territory.
A 2021-2025 project by AIMS and Parks Australia developed high-quality GIS datasets mapping emergent and shallow marine features in the Coral Sea Marine Park. Features mapped include coral atoll platforms, reef boundaries, depth contours, and cays, using composite satellite imagery from Sentinel 2, Landsat 8-9, and Sentinel 3, with manual mapping and review. The goal was to improve the precision and spatial detail of existing reef maps, providing permanent identifiers for features.
A 260 MB collection of styled geospatial datasets designed as a worldwide web mapping baselayer, with a focus on Australia and the Great Barrier Reef. The baselayer combines global Natural Earth 2 data with high-resolution Australian (100k) and Queensland (50k) coastlines, and includes country outlines, cities, and reef features. This dataset is filed in the eAtlas enduring data repository and has been superseded by a higher-resolution version.
NSW Government analysis of vehicle, person, and vessel trips at marinas, comparing new survey data with historical rates from 1978, 1990, and 2006-2008. The report, published by The Transport Planning Partnership, aims to update guidelines for marina developments across Australia. It includes surveys covering a range of marina sizes, locations, and facility types.
Regional Trends Online Tables is a regular source of official statistics for the Statistical Regions of the UK, produced by the Office for National Statistics. It includes a wide range of demographic, social, industrial, and economic statistics covering aspects of life in the regions. The dataset is designated as National Statistics and was last updated in July 2026.
Local Enterprise Partnership Profiles aim to help LEPs use official statistics to understand the economic, social, and environmental picture for their covered local authority areas. The dataset is produced by the UK Office for National Statistics and is designated as supporting material. It was last updated on 2026-07-08.
Road distances to nearest General Practice (GP) premises, primary schools, post offices, and supermarket/convenience stores for Lower Layer Super Output Areas (LSOAs) in England. The data is from 2004, using source data from 2001 to 2003, and was published by Neighbourhood Statistics from the Office of the Deputy Prime Minister.
England's road distances from Lower Layer Super Output Areas (LSOAs) to nearest General Practice premises, primary schools, post offices, and supermarkets in 2007. The dataset is administrative data sourced from Communities and Local Government (CLG) and published by Neighbourhood Statistics via the Office for National Statistics.
ID 2004 is a socioeconomic indicator measuring barriers to housing, specifically the difficulty of accessing owner-occupation based on house prices relative to incomes. The dataset, published by Neighbourhood Statistics and sourced from the Office of the Deputy Prime Minister, covers Local Authority Districts in England for the years 2002 and 2004. It is part of the Wider Barriers Subdomain and consists of administrative data with statistical transformations applied.
An administrative dataset from the UK's Office for National Statistics measuring barriers to housing services. The data provides a Difficulty of Access to Owner-Occupation indicator, likely based on house price to income ratios, for Local Authority Districts in England. It was published by Neighbourhood Statistics as part of the Wider Barriers Subdomain in 2007.