Loading...
Loading...
Text classification, translation, QA, summarization, dialogue, sentiment analysis, language modeling, text corpora
49,576 datasets
PAN20 Authorship Analysis data supports the detection of author changes in multi-author documents at the paragraph level. The dataset includes two topical subsets, one narrow (technology) and one wide (adding travel, philosophy, economics, history), each split into training, validation, and test sets. The data was provided by Eva Zangerle of Universität Innsbruck for the PAN 2020 shared task.
Three sediment cores from Nara Inlet, central Great Barrier Reef, covering the last 3000 years. The data, hosted by the Australian Ocean Data Network, describes sediment accumulation rates and composition changes in a tropical mixed clastic/carbonate system. It shows a decrease in both clastic and carbonate accumulation over time, with clastic input decreasing faster.
Oscar Douglas Skelton, a prime architect of early Canadian foreign policy, is honored by this lecture series inaugurated in 1991. The series features distinguished speakers examining topics related to Canada's foreign policy, international development, and trade. It is coordinated by Global Affairs Canada’s Open Insights Hub, which produces analysis on strategic issues and engages with external experts.
200 unique simulated particle decay topologies generated using the PhaseSpace library, with 16,000 samples per topology. The dataset provides leaf node particle features as input and Lowest Common Ancestor Generations (LCAGs) as training targets for the paper 'Learning Tree Structures from Leaves For Particle Decay Reconstruction'. It was created by J. Kahn at the Karlsruhe Institute of Technology.
Approximately 2800 computationally generated decoy structures exist for each of 49 antibody Fv domains, created using the RosettaAntibody protocol with stringent template exclusion criteria. This dataset provides a benchmark for evaluating scoring functions by their ability to identify near-native conformations from decoys based on provided H3-loop RMSD values. The structures originate from a curated set first compiled by Marze et al. in 2016 and were used in prior studies of CDR-H3 loop modeling.
Digital vector boundaries for Fire and Rescue Authorities in England and Wales, as of December 2024. The data provides full-resolution boundaries, typically following the Mean Low Water mark but sometimes extending to include offshore islands. It contains intellectual property rights from both Ordnance Survey and the Office for National Statistics.
ChinaHighO3 is a long-term, high-resolution dataset of ground-level maximum 8-hour average (MDA8) ozone concentrations across China. It was generated from big data sources using artificial intelligence by Jing Wei at the University of Maryland, College Park, and covers the period from 2013 to 2020. The dataset has a reported daily cross-validation R² of 0.87 and an RMSE of 17.10 µg m⁻³.
A post-processing scheme for recovering time correlation properties from thermostatted trajectories in quantum molecular dynamics simulations. The approach, introduced by Venkat Kapil of École Polytechnique Fédérale de Lausanne, yields spectroscopic observables with accuracy comparable to more demanding full path integral techniques. The dataset likely contains simulation results for model and realistic molecular systems.
1989-1997 spectral metocean data for the North Sea, including Significant Wave Height (HSIGN) and wave energy period (TMM10). The dataset was produced by Dr George Lavidas of Delft University of Technology as part of the WAVREP project funded by the European Union's Horizon 2020 programme. It is accompanied by publications detailing its calibration, validation, and production.
A 38-year database of spectral metocean conditions for the North Sea, covering 1998 to 2004. It contains yearly timesteps for Significant Wave Height (HSIGN) and wave energy period (TMM10) at a spatial resolution of 0.025 degrees. The dataset was produced by Dr. George Lavidas as part of the WAVREP project funded by the European Union's Horizon 2020 programme.
The North Sea Wave Database (NSWD) contains spectral metocean condition data for the North Sea region. It includes yearly timesteps for Significant Wave Height (HSIGN) and wave energy period (TMM10), with a spatial resolution of 0.025 degrees. The dataset was produced by Dr George Lavidas of Delft University of Technology as part of the WAVREP project funded by the European Union's Horizon 2020 programme.
Six years of spectral metocean data for the North Sea, covering 2012 to 2017. The dataset contains Significant Wave Height (HSIGN) and wave energy period (TMM10) measurements in meters and seconds. It was produced by Dr. George Lavidas at Delft University of Technology under the WAVREP project, funded by the European Union's Horizon 2020 programme.
2,340 test cases generated by EVOSUITE for 100 Java classes form a manually labeled dataset for six types of test smells. This dataset was created to benchmark the performance of test smell detection tools on automatically generated test suites. The analysis reveals that existing detection strategies misclassified over 70% of test smells in this context, highlighting patterns for tool improvement.
NSTA Infrastructure Data represents physical structures and facilities installed on the UK Continental Shelf, reported to the North Sea Transition Authority under the 2016 Energy Act. The dataset includes removed surface installations, subsea infrastructure, pipelines, and pipeline freespans, and is updated on a six-monthly cycle in April and October. Data is provided by the North Sea Transition Authority and was last updated on 2026-07-21.
Post-16 Education: Learner Participation and Outcomes in England is a National Statistics release from the UK Department for Education. The dataset, sourced from the Business, Innovation and Skills agency, focuses on further education and skills for learners aged 19 and over. It was last updated on July 8, 2026.
Waste heat sources in Vienna are analyzed for their potential to supply neighboring buildings. The dataset is published under a CC-BY-4.0 license by a cooperation of OGD Österreich and Wikimedia Österreich. Available file formats include CSV, JSON, and multiple geospatial formats like KML and ESRI SHAPE.
More than 70 mineral commodities, including economically important metals and materials, are tracked with annual production statistics by mass for individual countries grouped by continent. Import and export statistics are also available, though only for years up to 2002. The data is compiled from primary, official sources and is used to support government policy, commercial strategy, and economic analysis.
A 2006 Privacy Impact Assessment (PIA) for the Export Controls Online System (EXCOL) by Global Affairs Canada. The PIA was required when the legacy paper-based permit system, operational since 1988, was redesigned as an electronic service. It documents the department's commitment to personal information protection and compliance with Management of Information Technology Security requirements.
A dataset of mobile app reviews manually annotated for sentiment based on appraisal theory concepts. The data was created by Ruping Zhang and last updated on June 1, 2026. It is a small dataset of 9.5 KB, stored in an XLS file.
A manually annotated dataset for sentiment analysis of mobile app reviews, created by Ruping Zhang and last updated in June 2026. The dataset contains reviews labeled as High or Low based on performance parameters aligned with appraisal theory. It is a small dataset, 5.5 KB in size, and is available in XLS format.