Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
1046 ceramic vessels from 44 archaeological features represent the largest and temporally most highly resolved collection of morphological pottery data for the Central European Neolithic. The dataset was compiled by the University of Bern's MET-project between 2014 and 2018. It includes a spreadsheet of nominal and numeric morphological variables, vessel silhouettes, and typological drawings.
Azeem M. Shaikh from the University of Chicago proposes a methodology for inference in observational studies with time-varying treatment adoption. The analysis assumes a Cox proportional hazards model for treatment timing and studies randomization tests for a null hypothesis of no treatment effect. The methodology includes a simulation study and an empirical application using synthetic control-based test statistics and tobacco legislation data.
60 WhatsApp chat sessions collected by master students at Radboud University Nijmegen between 2013 and 2015. All participants are over 18 and provided consent, with conversations anonymized for research. The corpus was archived in 2018 for the CLARIAH-sponsored ACAD project and includes metadata in CMDI XML files.
Radboud University Nijmegen researchers collected 24,441 hand-drawn icons for crisis management systems from 32 Dutch student volunteers aged 19-30. The NicIcon database provides both offline scans and online pen trajectory time series for 14 classes of symbolic gestures representing events like floods and accidents. Each participant also wrote the 'London Letter' for forensic handwriting analysis.
MELTS software input and output files for modeling melt crystallization in the Martian lithosphere. The data, generated by J. Schools at the University of Maryland, College Park, is organized by oxygen fugacity, water content, lithosphere thickness, and mantle potential temperature. Associated MATLAB functions for analysis are available on a linked GitHub repository.
Fossil fuel COโ emission fluxes are provided on an hourly, gridded basis for atmospheric COโ modeling, specifically for the OCO2 Model Intercomparison Project. The dataset synthesizes monthly ODIAC inventory data with daily Carbon Monitor observations, using scaling factors to extend emissions from 2000 through at least 2025. It includes derived hourly global totals for verification and incorporates Carbon Monitor source data as NetCDF files for recent years.
595 female patients with type 2 diabetes from a hospital in Chengdu, Sichuan, between November 2018 and April 2023. This retrospective study by Yuqin Gan identifies nonlinear associations and specific thresholds between fasting insulin, visceral fat area, and estimated glomerular filtration rate. The analysis provides quantitative evidence for renal function risk stratification in this patient population.
A dataset of 200 bulk crystalline materials spanning a wide structural and chemical space, used to test an automated workflow for generating Maximally-localised Wannier functions (MLWFs). The workflow, implemented in an AiiDA environment by Valerio Vitale of the University of Cambridge, assesses MLWF quality by comparing band-structure interpolation accuracy against full first-principles calculations.
Rens van de Schoot's dataset from Utrecht University is used for online statistics training. It contains data from a 2010 study investigating the association between popularity status and antisocial behavior among 1,491 at-risk adolescents, with gender and ethnic background as moderators. The dataset includes variables for respondents' number, ethnic background, gender, socially desirable answering patterns, and covert and overt antisocial behavior.
Belgian day-ahead electricity market data used to model the societal effects of large-scale battery energy storage systems. The dataset includes modelling assumptions and input data from public sources like the ENTSO-E Transparency Platform and Elia's grid data. It was created by Emilia Rocha Ojeda of Utrecht University to accompany a conference paper.
Utrecht University researcher Thomas Smits compiled this dataset for the article 'Distant reading patterns of iconicity in 940.000 online circulations of 26 iconic photographs'. The material includes tabular metadata on webpages hosting reproductions of 26 iconic photographs, identified via the Google Cloud Vision API, and associated Doc2Vec text embeddings. The dataset supports analysis of how iconic images circulate and are discussed online.
A thesis by Milton Huang of the University of Michigan reexamines the concept of presence in Virtual Reality through the lens of emotional engagement. The work argues that emotions are essential to understanding presence and discusses validated psychological techniques for assessment. The dataset likely contains the textual content of this academic paper.
Three mapped stream networks from the upper Studibach catchment in Alptal, Switzerland, for dates in 2016 and 2018 representing extremely dry, dry, and wetting-up conditions. The data was created for a hydrological study published in Hydrology and Earth System Sciences by researchers including Rick Assendelft from the University of Zurich. It includes shapefiles for the mapped networks, a topographic network, and catchment boundaries.
Late Archean-Proterozoic geological features characterize the Harris Greenstone Domain, an arcuate tectonostratigraphic terrane in the centre of the Gawler Craton in Australia. The dataset is a georeferenced GeoPDF map depicting the domain's structure, including the Harris Greenstone Belt with its komatiite sequences and potential for Ni-Cu-PGE sulphide and lode-Au mineralising systems. The map's interpretation is based on aeromagnetic, gravity, and diamond drillcore data, with layers that can be toggled for customized analysis.
Four waves of semi-structured interviews with parents who experienced a birth in the year 2000, conducted by Paula England of Stanford University. Both mothers and fathers participated in individual and couple interviews between 2000 and 2005. The interviews covered views on parenthood, child-rearing responsibilities, household income, and the relationship between parents.
S66-BSIE contains energies and geometries for molecular complexes from the S66x8 and S66x100 datasets. The S66x100 dataset provides one hundred geometries per complex across a range of separation distances, with energies computed using the Psi4 package at RHF or B3LYP theory levels with multiple basis sets. The dataset was created by Sรธren Holm of Stanford University, extending the original S66x8 work by Goerigk, Kruse, and Grimme.
Twelve mainstream speech generation techniques were used to create fake audios for this dataset. It contains two versions, clean and noisy, with the noisy version created by adding noise from three databases at five different signal-to-noise ratios. The dataset is structured into training, development, and test sets, with a further split into seen and unseen subsets to evaluate model generalization.
Forestry Commission Scotland administrative boundaries are available as an interactive web map service. The service includes layers for FC Conservancy boundaries, FC Forest District boundaries, Woodlands In & Around Towns (WIAT), and the Central Scotland Green Network boundary (CSGN). Forestry Commission Scotland hosts the service with cooperation from Scottish Natural Heritage.
Euro-Calliope provides ready-to-use, open-source models of the European electricity system built with the Calliope framework. The models are available at three spatial resolutions: continental (1 node), national (34 nodes), and regional (497 nodes). Each model is a linear optimization problem that minimizes the total monetary cost of building renewable generation, balancing, and storage capacities to meet historically-based electricity demand across interconnected locations.
A dataset of probabilistic storm surge water level estimates for the Bengal delta coastline, spanning Bangladesh and India, at return periods from 25 to 500 years. It was generated by Md Jamal Uddin Khan of the Centre National de la Recherche Scientifique using a coupled hydrodynamic model forced by an ensemble of approximately 3600 synthetic cyclones. The data provides high-resolution (250 meters at the coast) estimates of maximum water elevation from dynamically modeled tide, surge, and wave interactions.