Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
tw-OCR is a dataset for optical character recognition tasks focused on Traditional Chinese. It contains images of real documents from Taiwan, such as forms, announcements, official documents, and academic materials, with annotations. The dataset was curated by Huang Liang Hsun and shared by Twinkle AI under a CC BY-SA 4.0 license.
21 dynamic scenes captured by 53 synchronized cameras, providing high-resolution multi-view video and 3D reconstructions. The dataset includes diverse human activities and object interactions recorded from a 360-degree perspective for volumetric video research.
DIS5K is a dataset for dichotomous image segmentation, ported from the DIS project. The dataset was created by Xuebin Qin and colleagues, with research presented at ECCV in 2022. Its specific scale and file formats are not detailed in the provided metadata.
Historical archival records from the Hechang Firm in Nagasaki contain annotated images of Suzhou numerals (0-9). The dataset documents trade activities between China, Japan, and Southeast Asia from 1880 to 1930. It was created by Morris0401 and is hosted on Hugging Face.
DynPose-100K provides 100,000 dynamic camera pose annotations, released in 2025 by researchers from NVIDIA, the University of Michigan, and NYU. The dataset focuses on camera trajectory estimation in video sequences and is accompanied by the Lightspeed benchmark for ground-truth validation.
H2R-1M is a family of large-scale robot-centric video datasets generated via a visual data augmentation pipeline. The dataset transforms human egocentric manipulation videos into robot-augmented versions by replacing human hands with simulated robot arms using pose estimation and physically aligned rendering. It was created by author yaoxu789 and last updated on Hugging Face on 2025-05-24.
A dataset last updated on 2025-05-30 contains reports on the consumption of communal resources by the executive committee of the Saksagansky district council in a city. The data is provided by the States site of Ukraine and is structured as separate reports for each resource, available in Excel XLSX format.
FlintstonesSV++ is an enhanced version of a dataset for story narration using visual scene graphs. It was created by researchers from the Insight Research Ireland Center for Data Analytics at the University of Galway. The associated paper was accepted at the Text2Story Workshop, ECIR Conference 2025 in Lucca, Italy.
A curated image classification dataset for adult content recognition across five distinct categories. It facilitates training models to distinguish explicit and non-explicit content in artistic, animated, and real-world imagery. The dataset was created by strangerguardhf and was last updated in May 2025.
2.3 million geo-referenced wildfire records represent 180 million acres burned across the United States from 1992 to 2020. The database, a sixth edition supporting the Fire Program Analysis system, includes core data elements like discovery date, final fire size, and precise point locations.
Ukraine's communal property land registry contains records for individual land plots. The data includes cadastral numbers, area, purpose codes, monetary valuations, and associated organizational identifiers. It was published by the States site of Ukraine and last updated on 2025-05-13.
A video dataset for evaluating action recognition models, compiled by author lixinhao and last updated on 2025-05-21. It includes videos from the ARID (Action Recognition in the Dark) and Breakfast (Action Recognition in Long Video) datasets. The dataset page provides links to download the original video files and annotations.
6,712 high-resolution portrait images of male and female individuals, each standardized to 1024×1024 pixels. The dataset was created by prithivMLmods and was last updated on 2025-05-19.
Over 3.7 million human preference responses evaluating AI-generated images, collected by Rapidata via its Python API. The data builds upon Google's research on Rich Human Feedback for Text-to-Image Generation and a previous smaller version. Collection involved 307,415 individual humans and was completed in less than two weeks.
visionxiang maintained this curated repository of aerial object detection resources, with the most recent update in July 2025. It aggregates links to specialized datasets, research papers, and code implementations for remote sensing tasks.
The dataset from adi2606 contains 1,941 anime character images organized into 322 folders, each representing a distinct anime series. It is intended for tasks like image classification and was last updated on April 29, 2025.
UrbanSyn provides over 7,500 photorealistic synthetic images of driving scenes. The dataset was created by the UrbanSyn organization to address the synth-to-real domain gap and includes ground-truth annotations for multiple vision tasks. It was last updated on the Hugging Face platform in April 2025.
A collection of datasets for training YOLOv8 object detection models. It includes annotated data for faces, hands, persons, and clothing items. The dataset was created by Blankse and last updated on Hugging Face in May 2025.
Deepfake-vs-Real-1440px-Max is a curated dataset of 28,000 portrait images for binary classification tasks. It was created by prithivMLmods and last updated on April 28, 2025. The dataset is intended for training and evaluating models in deepfake detection and media authenticity analysis.
A repository consolidated on May 19, 2025 by kamruzzaman-asif, containing high-quality instruction-tuning data from multiple sources. The dataset is structured for training and evaluating instruction-following models and includes splits like 'OdiaGenAI', which is described as a large-scale collection of diverse Bangla instructions and responses.