Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,653 datasets
Kinetics-700 is a large-scale collection of YouTube video URLs curated for human action recognition. The dataset was uploaded by atalaydenknalbant and was last updated on the Hugging Face platform in August 2025. It consists of 22 compressed archives that must be downloaded and decompressed to access the complete video collection.
Ultralytics COCO8-pose is a small dataset composed of the first 8 images from the COCO train 2017 set, split into 4 training and 4 validation images. It is designed for testing and debugging object detection models or experimenting with new detection approaches. The dataset was created by Ultralytics and was last updated on 2025-08-03.
InternRobotics released InternScenes in 2025, a large-scale collection of interactive indoor scenes featuring realistic layouts for embodied AI research. Presented at NeurIPS 2025, the data focuses on scene generation and robotic interaction within complex 3D environments.
Oregon workers' compensation claims counts are provided by the state's Department of Consumer and Business Services. The data covers claims where available since 1968, the year Oregon's modern workers' compensation system began. The dataset includes categories such as Permanent Disability and Fatalities.
LUMA is a multimodal dataset including audio, text, and image modalities, intended for benchmarking multimodal learning and multimodal uncertainty quantification. The dataset was authored by bezirganyan and last updated on 2025 -08-14. The full description and code are available on the dataset's Hugging Face page.
A dataset constructed by filtering cybersecurity-related text from FineWeb, a refined version of Common Crawl. It was created by Trend Micro's AI Lab and last updated on August 9, 2025. The dataset uses a high-quality seed set of manually curated cybersecurity text to train a binary classifier for filtering.
AncientDoc is a benchmark dataset for Chinese ancient document understanding created by yuchuan123 and hosted on Hugging Face. It contains 2,973 pages from approximately 100 documents across 14 literary types, spanning from the Warring States period to the Qing dynasty. The dataset supports multiple tasks including page-level OCR, vernacular translation, and reasoning-based question answering.
An image dataset for classifying Not Safe For Work content, uploaded by DarkyMan to Hugging Face. The dataset was last updated on September 11, 2025. Its specific size, format, and annotation details are not provided in the available metadata.
A labelled split for the Global Wheat Full Semantic Organ Segmentation (GWFSS) dataset, associated with a paper published in 2025. The dataset is hosted on HuggingFace by the author 'GlobalWheat' and was last updated on August 22, 2025. An unlabelled split and a benchmark model are also available via provided links.
DogSpeak is a large-scale, "in-the-wild" canine vocalization dataset sourced from tens of thousands of online social media videos. It was created by ArlingtonCL2 and last updated on August 14, 2025. The dataset captures a wide array of natural, organic interactions to advance research in animal communication and computational bioacoustics.
9,660 high-resolution RGB images categorized for detecting urban infrastructure problems. The dataset, created by Programmer-RD-AI and last updated in August 2025, focuses on issues like potholes, damaged roads, broken signs, illegal parking, and environmental cleanliness for computer vision applications.
Flemish vocabulary dataset organized by linguistic categories such as pronouns, verbs, nouns, and adjectives. It contains over 200 core Flemish words and includes 280,000 synthetic examples, 160,000 pattern examples, and other data types like conversation and translation. The dataset was created by author 0xnu and last updated on 2025-08-15.
The Superintendent of Documents (SuDocs) classification guidelines detail policies for assigning classification numbers to all U.S. Government publications. Published by the Government Publishing Office, this resource conveys the official system for organizing Federal Government publications regardless of format.
44,387 instances of question-to-Cypher query pairs, curated by Neo4j's GenAI team. The dataset is split into 39,554 training and 4,833 testing examples, each containing a question, schema, and Cypher query triplet. It was published on HuggingFace and updated in August 2025.
An annotated synthetic dataset of 500 Piping and Instrumentation Diagrams (P&IDs) incorporating different types of noise and complex symbols. The dataset, Digitize-PID, was created by Paliwal, S., Jain, A., Sharma, M., & Vig, L. and published in 2021. It contains only symbols, formatted for object detection tasks.
Redteam-Operations-Datasets provides structured records for offensive security operations in Active Directory and hybrid environments. The dataset is organized by attack phase, tactic, and technique, with each record following a canonical JSON schema. Created by elementalsouls and updated on August 17, 2025.
5 distinct image sources across real and synthetic categories provide training data for forgery detection, specifically utilizing DiffusionDB and LAION-Aesthetics. Evaluation sets derived from Midjourney, PixArt-alpha, and GPT-4o allow for testing cross-generator generalization.
STRIDE contains approximately 82 billion tokens arranged into 6 million visual sequences derived from 131,000 panoramic road images. Developed by Tera-AI, this dataset combines imagery with metadata and highway system data to enable generative world model training. The dataset was last updated in August 2025.
Personal Finance Reasoning-V2 is a dataset that won first prize in the Reasoning Datasets Competition organized by Bespoke Labs, HuggingFace & Together.AI in April-May 2025. It focuses on personal finance, contrasting with benchmarks for corporate finance and algorithmic trading. The dataset was created by Akhil-Theerthala and was last updated on August 11, 2025.
A large-scale multilingual document OCR dataset containing approximately 400GB of images with annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. It was created by Nayana-cognitivelab and last updated on 2025-07-21.