Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,658 datasets
STRIDE contains approximately 82 billion tokens arranged into 6 million visual sequences derived from 131,000 panoramic road images. Developed by Tera-AI, this dataset combines imagery with metadata and highway system data to enable generative world model training. The dataset was last updated in August 2025.
Personal Finance Reasoning-V2 is a dataset that won first prize in the Reasoning Datasets Competition organized by Bespoke Labs, HuggingFace & Together.AI in April-May 2025. It focuses on personal finance, contrasting with benchmarks for corporate finance and algorithmic trading. The dataset was created by Akhil-Theerthala and was last updated on August 11, 2025.
A large-scale multilingual document OCR dataset containing approximately 400GB of images with annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. It was created by Nayana-cognitivelab and last updated on 2025-07-21.
MAESTRO (MIDI and Audio Edited for Synchronous Tracks and Organization) v3.0.0 is a dataset of about 200 hours of virtuosic piano performances. It features fine alignment, approximately 3 milliseconds, between note labels and audio waveforms. The dataset was created by projectlosangeles and mirrored on Hugging Face as of August 3, 2025.
Astra77's Huaxia Lib aggregates a variety of ancient Chinese texts primarily sourced from the 殆知閣 (Almost Know Pavilion) repository. The collection has undergone organization and basic data cleaning and is described as one of the most comprehensive classical Chinese datasets available online. It was last updated on August 8, 2025.
The dataset contains information about the organizational structure of the municipal enterprise "City Ritual Service" in accordance with the organizational and administrative document "Statistical painting". It was published on the States site of Ukraine and last updated on August 5, 2025.
A dataset from the Open Paws initiative, last updated on August 6, 2025, designed to train AI systems on animal welfare, rights, and liberation principles. It is formatted as JSONL and is multilingual, with a primary focus on English. The dataset is intended to support the development of AI that understands and promotes ethical reasoning related to animal advocacy.
Digitize-PID is an annotated synthetic dataset of 500 Piping and Instrumentation Diagrams (P&IDs) containing only symbols for object detection tasks. The dataset incorporates different types of noise and complex symbols. It was created by Paliwal, S., Jain, A., Sharma, M., & Vig, L. (2021) and is hosted on Hugging Face by user hamzas.
A Web Map Service (WMS) layer provides the urban development plan 'Europastraße - Gant (origin)' for the city of Nürtingen. The plan is described as final and original, sourced from the XPlanung 5.0 standard. The service is provided by the Bundesamt für Kartographie und Geodäsie and was last updated on August 14, 2025.
HotpotQA data is enriched with synthetic long-context examples to push the boundaries of multi-hop reasoning. The dataset, provided by BytedTsinghua-SIA, is designed for training and evaluating long-context language models like MemAgent. It was last updated on July 30, 2025.
An example dataset specifying the required format for finetuning the Surya optical character recognition model. It was created by datalab-to and last updated on August 8, 2025. The dataset description outlines the need for paired image and text columns, with specific markup for mathematical content.
U.S. Census Bureau's 2023 TIGER/Line shapefile provides geographic data for Michigan's primary and secondary roads. The dataset includes primary roads, which are divided, limited-access highways, and secondary roads, which are main arteries in the U.S., State, or County Highway systems. Each road type is identified by a specific MAF/TIGER Feature Classification Code (MTFCC).
Mobius Data provides DNA methylation data from blood samples of individuals with Myalgic Encephalomyelitis/Chronic Fatigue Syndrome (ME/CFS), Long COVID, and healthy controls. The dataset, created by VerisimilitudeX and last updated on August 9, 2025, is derived from Illumina HumanMethylation450 BeadChip and MethylationEPIC arrays. It is organized to facilitate research into epigenetic biomarkers for these post-viral illnesses.
99 annotated flowchart images provide a specialized computer vision dataset for detecting structural elements. FC-Detection contains detailed bounding box annotations for 9 different flowchart components. The dataset was created by Galirage Inc. and was last updated on the Hugging Face platform in July 2025.
Synthetic OCR dataset of Russian technical drawings text generated with TextRecognitionDataGenerator. Texts include engineering terms, GOST standards, dimensions, specifications, and full Cyrillic alphabet coverage. The dataset was created by Mkz-Prog and last updated on Hugging Face in August 2025.
70,000 text prompts and 68,000 manually annotated images organized into a hierarchy of 12 tasks and 44 categories. The collection covers safety domains including toxicity, fairness, bias, and privacy to evaluate text-to-image generation models.
54,120 frames of multi-sensor data including camera, radar, LiDAR, and GPS/IMU for inland water surface perception. The dataset provides annotations for object detection, semantic segmentation, and instance segmentation across 7 categories such as boats, humans, and buoys.
yolay created this dataset for the official implementation of the paper "Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models". The dataset addresses challenges LLMs face with complex instructions containing multiple constraints in parallel, chained, and branching structures. It was last updated on 2025-07-31.
SAMHSA's Mental Health Treatment Facilities Locator is an online resource listing facilities providing services to persons with mental illness. The data is updated annually via the National Mental Health Services Survey, with the most recent complete update including data collected as of April 30, 2010. New facilities are added monthly, and details like names and addresses are updated weekly based on facility reports.
The dataset from the States site of Ukraine contains information about fairs within the Vladimir City Council territory. It includes identifiers, fair names, rental costs, addresses, duration, schedules, organizers, and their contact details. The dataset was last updated on 2025-08-06.