Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
ERQA is an evaluation benchmark adapted from embodiedreasoning/ERQA, focusing on real-world scenarios for robotics. The dataset, hosted by FlagEval, was last updated on April 22, 2025, and covers topics related to spatial reasoning and world knowledge. It was originally provided in TFRecord format and has been converted for easier use.
A multimodal dataset for training optical character recognition models on the Khmer language, created by KiteAether and hosted on Hugging Face. It was last updated on June 11, 2025. The dataset is categorized as containing between 1,000 and 10,000 samples.
UGround V1 Data is a computer vision dataset from the osunlp organization, last updated on May -1, 2025. The description notes an addition of a bounding box version of 'Web-Hybrid' data, with normalized coordinates. Access to the dataset requires an application process involving contact via email.
Presenting a sample of 4,134 individuals for infrared face recognition. Data was collected using a RealSense D453i device across indoor and outdoor scenes, featuring subjects from children to the elderly, with young and middle-aged adults as the majority.
Comprising handwriting OCR data from 100 subjects, including 50 Japanese, 49 Korean, and 1 Afghan individual. It was collected using multiple cellphone models and features different text corpora per subject.
Three Excel resources list the Department of Inspection Work, its structural subdivisions, and officials for Sumy City Council. The directory includes identification codes from the Ukrainian Unified State Register, email addresses, phone numbers, and physical addresses. This data was last updated on the States site of Ukraine on April 25, 2025.
DataSalon hosts a dataset from the State site of Ukraine on buildings with implemented energy management systems. The dataset includes buildings of state, private, and communal property. It was last updated on April 28, 2025.
StrongBaseline YOLOv12-BoT-SORT-ReID is a dataset of video demonstrations for object detection and tracking models. The dataset, created by author wish44165, was last updated on April 30, 2025. It contains preview and full videos showing model inferences and single-frame enhancements from a CVPR 2025 workshop.
Google Scanned Objects Symmetry Axis Dataset (GSO-SAD) is an extension of the Google Scanned Objects dataset, enriched with symmetry axis annotations for each 3D scanned object. The dataset is designed to assist in pose estimation tasks by providing explicit symmetry information for objects with geometric and texture symmetries. It was created by SEU-WYL and last updated on Hugging Face in April 2025.
High-resolution images of Devanagari script characters address limitations of existing low-resolution datasets. The collection includes vowels, consonants, matra combinations, and Hindi numerals. Author Mayank022 uploaded this dataset to Hugging Face on April 21, 2025.
DEArt is a dataset of European paintings from the 12th to the 18th centuries. It contains more than 15000 images, with manual annotations for bounding boxes identifying 69 object classes and 12 possible poses for human-like objects. The dataset was created by author 'biglam' and last updated on March 31, 2025.
A dataset published on 2025-04-30 by the States site of Ukraine. It contains information about the organizational structure of an information manager, as defined by an organizational and administrative document. The data is provided in an EXCEL XLSX file format.
FlagEval provides a benchmark for evaluating embodied spatial understanding in large vision-language models. The dataset contains 3,640 question-answer pairs automatically derived from embodied scenes, covering six spatial relationships from an egocentric perspective. It was adapted from an original image-format dataset by Phineas476 and last updated on 2025-04-21.
RuixuanJiang's repository hosts four low-level visual datasets. The collection was last updated on May 23, 2025. Specific details on the individual datasets' content and scale are not provided in the metadata.
100,000 synthetic arithmetic problems categorized into addition, subtraction, multiplication, and division. The collection includes 75,000 numeric problems evenly distributed across four basic operations with 18,750 samples each.
Sample of 122 images for passenger behavior recognition. It includes images of multiple age groups and races (Caucasian, Black, Indian) depicting normal and abnormal behaviors such as carsick, sleepy, and lost items.
Over 6,000 images of female anatomy, intended for captioning or training captioning models. The dataset includes a list of filenames and a subdirectory containing text captions generated by joycaption, which share the same filename as the parent image. It was created by the author 'hardlyworking' and last updated on April 8, 2025.
The organizational structure of the Main Department of the Pension Fund of Ukraine in the Dnipropetrovsk region. The dataset was last updated on 2025.04.24 and is provided by the States site of Ukraine via the eu_open_data platform.
A collection of Traditional Chinese text for optical character recognition (OCR) tasks. It was created by esun-ai and last updated in June 2025. The specific number of rows, columns, and file formats are not provided.
Image Captioning Openimages Subset is a dataset for training and evaluating image captioning models. It was published on huggingface by user erayalp and was last updated on 2025-05-26. The dataset likely contains images paired with descriptive text, derived from the OpenImages repository.