Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,347 datasets
Netherlands-based survey of employer establishments on labor demand and personnel policy, conducted biennially since 1989. The dataset likely contains around 3,000 observations per wave, with 15 measurement waves available to researchers. It is managed by the Netherlands Institute for Social Research (SCP) and originated from the Organisation for Strategic Labour Market Research (OSA).
GRAID NuImages is a question-answer dataset generated by the GRAID framework for enhancing spatial reasoning in Vision-Language Models. The dataset was created by author kd7 and is associated with a research paper and project page. It was last updated on October 29, 2025.
The Startmonitor is a survey project by the Dutch Ministry of Education, Culture and Science (OCW) and ResearchNed. It tracks the information use, study choice process, and first-year integration and satisfaction of new students in higher education to identify determinants of study success and dropout. This deposit contains the September 2020 survey, part of the 2020-2021 academic year cycle, which has been conducted annually since 2008-2009.
1989 to 2016 panel survey of employer establishments in the Netherlands, conducted biennially. It contains around 3,000 observations per wave and 14 measurement waves, designed to provide insight into labor demand and personnel policy. The data is managed by the Netherlands Institute for Social Research (SCP) and originated from the Organization for Strategic Labor Market Research (OSA).
The Arbeidsvraagpanel (AVP) is a biennial survey of employer establishments in the Netherlands, established in 1989. The 2017 supplement, managed by the Sociaal en Cultureel Planbureau (SCP), provides additional data on employee counts, training days, and detailed sector codes. This dataset is part of a series with around 3,000 observations per wave and 15 measurement waves available to researchers.
Pre-COOL is a longitudinal study tracking two cohorts of children to understand the effects of different forms of childcare and early childhood education. The study, conducted by the Kohnstamm Instituut and the University of Amsterdam, collects data on children's cognitive and socio-emotional development, their family backgrounds, and the quality of preschool and kindergarten facilities. The dataset includes a four-year-old cohort started in 2009 and a two-year-old cohort started in 2010, with data collection continuing through the end of primary school.
The TotalSegmentator Organs dataset contains CT scans with dense segmentation annotations for 14 anatomical structures. The dataset is provided by MedOtter and was last updated on October 30, 2025. The data format is NIfTI (.nii.gz).
Code-170k-shona is a dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Shona. It was created by michsethowusu and last updated on October 30, 2025. The dataset aims to make coding education accessible to Shona speakers through multi-turn dialogues.
A subset of 100,000 anime character images from the Zerochan webdataset. The images are filtered to be non-monochrome and depict a single person, head, and face with one primary character. Annotator animetimm created this dataset, which was last updated on November 5, 2025.
OpenPecha's benchmark dataset evaluates Tibetan optical character recognition models. It includes diverse scripts, writing styles, and print methods to enable testing across multiple domains. The dataset was last updated on October 30, 2025.
GRAID NuImages was generated using the GRAID framework, which transforms object detection annotations into structured question-answer pairs. The dataset tests various aspects of object understanding and spatial reasoning. It was created by authors Charles Xu, Qiyang Li, Jianlan Luo, and Sergey Levine.
High-resolution images of handwritten mathematical notes in English, including problem statements, worked examples, formulas, and annotated derivations. The dataset was created by HumynLabs and was last updated on the platform on 2025-10-23.
Dhiraj45's dataset contains over 5,000 high-resolution anime-style images. The collection is described as being curated from Comix Wave Anime and is intended for machine learning and creative projects. It was last updated on the Hugging Face platform in October 2025.
Filled with survey results from employees in the public sector, conducted in 2022. It is a co-production of the Dutch Ministry of the Interior and Kingdom Relations and Statistics Netherlands (CBS). The data is documented with variable descriptions and supporting attachments.
An open-source synthetic dataset for computer vision object detection tasks. It contains high-quality, realistic images simulating people holding weapons in public areas from CCTV camera perspectives. The dataset was created by Simuletic and was last updated on October 28, 2025.
Over ten thousand pages of primary source material from the Kurdistan Workers' Party's monthly bulletin Serxwebun, spanning 1982 to 2015, were analyzed by researcher ΓaΔlayan BaΕer. The data repository contains representative excerpts focusing on women's roles, mobilization, and integration within the rebel organization. This analysis shows that women insurgents enable tactical diversity and aid the organization's survivability.
Sarah Bush's Annotation for Transparent Inquiry project explores the relationship between human rights and democracy promotion fields. The research relies on primary source material from state agencies, international organizations, philanthropies, NGOs, and 20 semi-structured interviews. It was harvested by the Qualitative Data Repository and last updated on October 20, 2025.
Marc Trachtenberg's appendices provide supplementary detail for his analysis of Cold War relations from 1945 to 1963. The eight appendices originated as footnotes for his book 'A Constructed Peace' and were moved online due to manuscript length constraints. They are archived on the Qualitative Data Repository (QDR) and use the same reference system as the published work.
Interview transcripts with NATO and EU officials conducted in Brussels and The Hague from 2016 to 2019. The collection focuses on organizational responses to cybersecurity, energy security, and peacekeeping threats. The research project, led by Ivan Dinev Ivanov, aims to develop a framework explaining how international organizations adapt to emerging challenges.
YouTube-GDD is a gun detection dataset sourced from YouTube videos, created by UCAS-GYX and updated in December 2025. It provides visual data for identifying firearms within diverse contextual environments to support deep learning and object detection research.