Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,416 datasets
Over ten thousand pages of primary source material from the Kurdistan Workers' Party's monthly bulletin Serxwebun, spanning 1982 to 2015, were analyzed by researcher Çağlayan Başer. The data repository contains representative excerpts focusing on women's roles, mobilization, and integration within the rebel organization. This analysis shows that women insurgents enable tactical diversity and aid the organization's survivability.
Sarah Bush's Annotation for Transparent Inquiry project explores the relationship between human rights and democracy promotion fields. The research relies on primary source material from state agencies, international organizations, philanthropies, NGOs, and 20 semi-structured interviews. It was harvested by the Qualitative Data Repository and last updated on October 20, 2025.
Marc Trachtenberg's appendices provide supplementary detail for his analysis of Cold War relations from 1945 to 1963. The eight appendices originated as footnotes for his book 'A Constructed Peace' and were moved online due to manuscript length constraints. They are archived on the Qualitative Data Repository (QDR) and use the same reference system as the published work.
Interview transcripts with NATO and EU officials conducted in Brussels and The Hague from 2016 to 2019. The collection focuses on organizational responses to cybersecurity, energy security, and peacekeeping threats. The research project, led by Ivan Dinev Ivanov, aims to develop a framework explaining how international organizations adapt to emerging challenges.
YouTube-GDD is a gun detection dataset sourced from YouTube videos, created by UCAS-GYX and updated in December 2025. It provides visual data for identifying firearms within diverse contextual environments to support deep learning and object detection research.
Federal Highway Administration data provides geographic boundaries for Metropolitan Planning Organizations (MPOs) across the United States. Compiled on November 10, 2025, this dataset supports multimodal transportation planning and programming. It includes planning locations, sizes, and names for metropolitan areas.
Esther Diaz Romanillos compiled this dataset through systematic documentary analysis of regional laws and official statistics from the Spanish Ministry of Education. It contains normative and quantitative data on rural schools in Spain, focusing on Colegios Rurales Agrupados (CRA) from the 2010/2011 to 2022/2023 academic years. The data includes historical series on schools, units, and students, disaggregated by autonomous community and year.
Causal2Needles is a benchmark dataset containing between 1,000 and 10,000 records designed for evaluating long-video understanding in multimodal large language models. Developed for the NeurIPS 2025 Datasets and Benchmarks Track, it focuses on "2-needle" reasoning tasks where models must retrieve and synthesize two separate temporal segments from a single video. The dataset provides multiple-choice questions that test causal reasoning across long-form visual content.
2018 data from the VSOI database, which contains information on research projects, scientific organizations, and researchers including professors and associate professors in the Netherlands. The data was used for the NARCIS portal and the former Dutch Research Database (NOD). The dataset is part of an archive series spanning from 2002 to 2020, with this specific entry covering the year 2018.
Wageningen, Netherlands survey from 1970 on church attendance, membership, and religious opinions. The data, authored by B. Wemmenhove and hosted by DANS, includes background variables on occupation, education, politics, and media exposure. It was last updated on the platform in October 2025.
176,999 programming conversations translated into Somali, sourced from the glaiveai/glaive-code-assistant-v2 dataset. The dataset was created by michsethowusu and last updated on Hugging Face in October 2025. It aims to make coding education accessible to Somali speakers.
2022-2023 data containing all IRS 990 organizations by Employer Identification Number (EIN) and their matched OpenCorporates unique identifiers. The dataset was created by GivingTuesday and allows linking nonprofit tax data with state-level corporate records. All 990-CN, EZ, and PF organization filings for the period were tested for matches.
A dataset of Spanish financial annual report paragraphs from 2014 to 2018, annotated for causality detection. The data was created for the FinCausal 2023 shared task, with linguists labeling cause and effect elements within each paragraph. It is published by e-cienciaDatos Harvested Dataverse and was last updated in October 2025.
CausalVerse Image Dataset contains two families of splits: Physics splits (Fall, Refraction, Slope, Spring) and Static image generation splits (scene1-4). The dataset is associated with the paper 'CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations' and was last updated on 2025-10-23.
IJCNN is a classic benchmark dataset from the 2001 International Joint Conference on Neural Networks competition. It contains data for a binary classification task, likely involving signal or pattern recognition. The dataset's origin and specific size are not detailed in the provided metadata.
176,999 programming conversations, originally from the glaive-code-assistant-v2 dataset, have been translated into the Kanuri language. The dataset was created by michsethowusu and last updated on October 30, 2025. It consists of multi-turn dialogues covering various programming concepts and algorithms.
Npix Open Image Dataset is hosted on Hugging Face by ArkAiLab-Adl. The dataset was last updated on December 13, 2025. Its specific content, scale, and annotations require verification after download.
Eleven interviews, nine audio and two video, capture the experiences of women who were members of the Dutch National Socialist Movement (NSB) between 1931 and 1945. The project was conducted by Aletta, Institute for Women's History in 2008 to document the life stories of women who chose the side of the occupier during World War II. The interviews focus on memories of NSB activities during the occupation, experiences after Dolle Dinsdag, and the period of internment.
A dataset from two studies examining how engaging leadership influences employee perceptions of organizational values, psychological need satisfaction, and work engagement. Study 1 used a cross-sectional design with 436 participants, and Study 2 used a longitudinal design across three time-points with 69 participants. The data was authored by L. van Tuin and deposited in the DANS Data Station Social Sciences and Humanities Collection on 2025-10-24.
Ongekend Bijzonder is an oral history project by Stichting BMP containing 248 in-depth, filmed interviews with refugees who arrived in the Netherlands over the past forty years. Each interview dataset includes a video, a transcription, and a description, with a focus on rebuilding lives and contributions to the four major Dutch cities. The project was conducted in 2017, with management of the interviews transferred to city archives.