Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,468 datasets
A.W.J. Bruekers published a dataset of passport copies from the Nederweert Magistrates' Court, covering 1773 to 1797. The register documents labor migrants, adventurers, and pilgrims, often including detailed physical descriptions and clothing of applicants. The data is hosted by the DANS Data Station Social Sciences and Humanities Collection and was last updated on October 24, 2025.
Ukrainian-language prompts and paired responses designed to counter Russian propaganda narratives. The dataset, created by lapa-llm, trains counter-narrative reasoning with fact-based replies grounded in international law and reputable sources. Its last update was recorded in November 2025.
250,000 images sampled from the ImageNet-21K dataset form this multimodal reasoning collection. For each image, the dataset provides a prompt and two different step-by-step reasoning tokens and outputs. Created by krishnateja95 and last updated on October 8, 2025, it is designed for research on multimodal summarization.
Approximately 1.04 million per-frame 2D observations of 557 globally unique static objects across 8 categories, captured on a university campus. The dataset was created by Dongmeyong Lee and last updated on October 15, 2025. It is intended for improving long-term robot perception and navigation under varying environmental conditions.
High-resolution images of English handwritten notes cover algorithm explanations, code snippets, flowcharts, and theoretical content. The dataset is authored by HumynLabs and was last updated on October 23, 2025. It is designed to support AI research in document understanding for computer science education.
UniqueData provides a collection of document images annotated for text detection. The dataset includes scanned papers, forms, invoices, and handwritten notes with varied layouts, fonts, and styles. It was last updated on October 10, 2025.
Qari Arabic OCR 10K is a dataset for optical character recognition of Arabic text, hosted on HuggingFace. The dataset was uploaded by author melsiddieg and was last updated on December 7, 2025. The specific content, scale, and collection methodology are not detailed in the available metadata.
The EPFL-Smart-Kitchen-30 dataset provides a benchmark for action recognition. It likely contains video clips of kitchen activities annotated with action labels. The dataset was created by amathislab and was last updated on October 23, 2025.
Stpls3D provides synthetic and real-world 2D and 3D point cloud data for semantic and instance segmentation, published by meidachen for BMVC 2022. The dataset utilizes AirSim for synthetic generation alongside real-world photogrammetry captures to support 3D reconstruction benchmarks.
Long-term Occupational Projections from data.ny.gov provide a 10-year forecast of job openings across New York State and its 10 labor market regions. The dataset includes annual estimates for total openings, growth openings, and replacement openings for detailed occupations. It was last updated in September 2025.
SPSLingAnc provides a mixed-methods analysis of complete fictional narratives in English, focusing on linguistic tokens that shape readers' conceptualization of storyworld possible selves. The dataset comprises 91,328 words from two novels, with 5,980 tagged tokens analyzed using MAXQDA 2022 software. It was authored by María-Ángeles Martínez and last updated on October 14, 2025.
17,091,657 face images of 360,232 distinct individuals comprise this collection for training face recognition models. The dataset was cleaned and released by yayoimizuha, with the latest update recorded in October 2025. It is designed to support state-of-the-art performance on benchmarks like IJB-C and MegaFace.
The Kinetics-700 dataset is a large-scale collection of YouTube video URLs for human action recognition, containing 700 distinct action classes. It is an extension of the original Kinetics dataset, with each video clip being approximately 10 seconds long and depicting a single human action. The dataset was uploaded by author 'bitmind' to the Hugging Face platform, with a last recorded update in October 2025.
60 research laboratories affiliated to the University of Lille contributed grey literature deposits to the French national open repository HAL. The dataset likely contains records for conference papers, reports, working papers, theses, and dissertations across STM and SSH fields. Author J. Schöpfel published this analysis via the DANS Data Station Social Sciences and Humanities Collection in October 2025.
High-resolution images of handwritten physics notes, including theoretical explanations, formulas, diagrams, derivations, and problem-solving steps. The dataset was created by HumynLabs and last updated on Hugging Face in October 2025. It is designed to support AI research in handwriting recognition and scientific document understanding.
PaddlePaddle published a demonstration dataset for visual-language optical character recognition on Hugging Face on 2025-11-27. The dataset likely contains examples for testing or demonstrating a PaddleOCR visual-language model. Its specific content, size, and structure are not detailed in the provided metadata.
270,000 PDF pages converted to plain-text using GPT-4.1 with a specialized prompting strategy to preserve natural reading order and born-digital content. Released by the Allen Institute for AI (allenai) in October 2025, this collection includes both the original PDF files and their corresponding high-fidelity OCR outputs. It serves as an improved successor to the olmOCR-mix-0225 dataset with cleaner processing and more consistent formatting.
1,000 document images with human-annotated quality scores and OCR text extracted by the Qwen2.5-VL-72B model. The dataset was created by Aslan-mingye and last updated on October 17, 2025. It is designed as a benchmark for evaluating OCR performance across diverse document types including academic papers, textbooks, and e-books in multiple languages.
Sixteen mudbrick samples from the site of Tell Zurghul/Nigin in Iraq, dating from the mid-5th to mid 3rd millennium BCE, were analyzed. The dataset contains raw data from empirical field tests assessing material composition, grain size, clay behavior, and cohesion. It was created by De Vito, Licia as part of the EnEAp project and last updated in October 2025.
36 items across 4 categories record Spanish R&D projects on media communication ethics from 2007 to 2022. The database maps university research productivity over 15 years, providing derived data from project analysis. It was authored by Polledo-Zulueta, Yenisley and last updated on October 14, -2025.