Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,665 datasets
552,992 high-resolution images categorized into 18 distinct hand gesture classes such as 'stop', 'peace', and 'fist'. Each image includes bounding box annotations for hand localization and was captured by 34,730 unique participants in diverse real-world environments.
Kamianka City Council's Department of Housing and Communal Services and Construction provides a directory of municipal enterprises under its management. The dataset likely contains identification codes, official websites, email addresses, telephone numbers, and physical addresses for these organizations. It was last updated on January 6, 2025.
Corrected versions of DermaMNIST, HAM10000, and Fitzpatrick17k dermatology datasets address data leakage and duplication issues identified in the original benchmarks. Created by kakumarabhishek and updated in February 2025, this collection provides refined data for skin disease classification. It focuses on improving the integrity of medical image computing benchmarks through rigorous quality analysis.
Published on HuggingFace by wubingheng on 2025-02-11. The dataset likely contains medical images for classification tasks. Columns suggest it is intended for training or evaluating vision models.
This repository, maintained by jasonmanesis and updated in January 2025, serves as a curated index of radar (SAR) and optical satellite datasets for maritime monitoring. It catalogs specific resources for ship detection, classification, and segmentation tasks, including benchmarks like xView, SSDD, and HRSID.
Measurements of potential microbial functions and soil properties in High-Arctic tundra biocrusts and mineral soils. Data was collected near Ny-Γ lesund, Svalbard, across seasons along a toposequence. The dataset is provided by ENVIDAT and was last updated in 2025.
A directory of enterprises, institutions, and organizations within the Shostka City Territorial Community in Ukraine, last updated on 2025-01-09. The dataset likely contains identification codes, official websites, email addresses, telephone numbers, and physical addresses. It is sourced from the Handbook of the State Registration Department of Shostka City Council and published on the eu_open_data platform.
A large-scale dataset curated for training unified image generation models. The dataset is associated with the OmniGen project and was last updated in December 2024. It was created by author yzwang to address the lack of readily available data for multi-task image processing.
CC-OCR is a benchmark dataset for evaluating large multimodal models in optical character recognition and literacy tasks. The dataset is hosted in a TSV format for use with the VLMEvalKit evaluation framework. It was created by author wulipc and last updated in December 2024.
TIGER-Lab's VISTA-400K is a synthetic dataset designed to enhance video understanding in large multimodal models. The data is generated by the VISTA pipeline, which creates long-duration, high-resolution video instruction-following examples. The dataset was last updated on the Hugging Face platform in December 2024.
VLGuard consists of image-text pairs categorized into training and testing sets for safety fine-tuning of Vision Large Language Models. The dataset includes metadata files in JSON format and corresponding image archives to facilitate the development of safety baselines for multimodal AI systems.
48,384 high-resolution aerial images containing 1,041,714 annotated instances across 18 object categories. The dataset focuses on oriented object detection (OBB) and was introduced in the IEEE TPAMI journal for remote sensing research.
Image segmentation, object detection, and keypoint labels are generated through this web-based framework created by jsbroks and last updated in January 2025. It facilitates the creation of structured computer vision datasets in the COCO format from user-provided image files.
A directory listing enterprises, institutions, and organizations managed by a specific information manager in Ukraine, likely including identification codes. The dataset is published on the eu_open_data platform by the States site of Ukraine and was last updated in January 2025. Its exact scope and size are not detailed in the provided metadata.
Information on the organizational and personnel work of the Executive Committee of Starokostiantyniv City Council, socio-economic development and socio-political situation in Starokostiantyniv. The dataset is published on the eu_open_data platform by the States site of Ukraine and was last updated on 2025-01-14. The raw description indicates it is a certificate, and the primary file format is WORD DOC.
BioTrove comprises a large collection of images for biodiversity research, curated by BGLab. The dataset includes well-processed metadata with full taxa information and URLs pointing to image files. The dataset page was last updated on December 13, 2024.
TurkishLLaVA OCR Enhancement Dataset is a specialized collection of 100,000 books sourced entirely from Turkish materials. Created by ytu-ce-cosmos, it was last updated on December 17, 2024. The dataset is designed to improve the Turkish Optical Character Recognition capabilities of the Turkish-LLaVA-v0.1 model.
92,421 lines of Chinese game dialogue and narration extracted from Honkai Impact 3rd gameplay videos, covering story chapters from 'Main Story 1: Dusk, Maiden, Warship' to 'Main Story Part 2 Chapter 03 Interlude: A Sleepwalker's Pain'. The dataset was created by author mrzjy using an AI pipeline involving video frame extraction, OCR, and visual language model processing, and was last updated on December 17, 2024.
600 PubChem organic molecules structured into 5 multimodal question-and-answering tasks for chemistry reasoning. The content focuses on atom counting for carbon and hydrogen and molecular weight calculations, divided into validation and evaluation subsets.
Infinity-MM provides between 10 million and 100 million multimodal instruction samples, developed by the Beijing Academy of Artificial Intelligence (BAAI) in 2024. It utilizes a synthetic generation pipeline to create detailed image annotations and diverse question-answer pairs for bilingual model training in English and Chinese.