Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Annual extract of select financial data items from IRS Form 990PF filings for private foundations. The dataset is produced by the Department of the Treasury for public release and was last updated in August 2024.
Focusing on COCO and YOLO label types, this resource provides a Python-based framework for managing computer vision bounding box annotations. It was developed by the pylabel-project and last updated in August 2024 to facilitate format translation.
MiniGPT-4 captions generated for ImageNet1k images. The dataset was created by author wufeim and was last updated on July 25, 2024. It can be used for training or fine-tuning diffusion models for image generation.
A speech dataset for the Shona language, authored by Beijuka and published on the Hugging Face platform. The dataset was last updated on August 8, 2024. Its specific content, size, and collection methodology are not detailed in the available metadata.
1.9 million data annotations in JSON format, created by OpenGVLab for instruction-tuning multimodal AI models. The dataset was last updated in June 2024. It is derived from image datasets like M3IT, which underwent quality filtering processes.
BDD100K-validation contains 10,000 samples from one of the largest open-source driving datasets. The dataset consists of images extracted every 10 seconds from driving videos, providing labels for object detection, weather, time of day, and scene. It was uploaded to Hugging Face by dgural and last updated on June 12, 2024.
The MSCOCO dataset contains over 330,000 images, with more than 200,000 labeled. It includes 1.5 million object instances across 80 object categories and 91 stuff categories for segmentation. Each image has 5 captions, and the dataset contains annotations for 250,000 people with keypoints.
A synthetic dataset of 300,000 images designed for optical character recognition tasks. It contains printed Cyrillic characters including the full Russian alphabet and common punctuation symbols. The dataset was created by author DonkeySmall and was last updated in July 2024.
A collection of documents related to the General Agreement on Tariffs and Trade (GATT), spanning over 50 years from January 1, 1946, to September 6, 1996. The dataset is organized into a single Parquet file containing document metadata and was uploaded by author PleIAs to Hugging Face on June 17, 2024.
New York City's Metropolitan Transportation Authority provides Wi-Fi and cellular service data for underground subway stations. The dataset includes station names, boroughs, lines, and service provider availability from a snapshot taken in 2015 and 2016. It was published by data.ny.gov and last updated in May 2024.
1.2 billion object instances across 3.5 million images featuring semantic tags, bounding boxes, and detailed captions. This collection facilitates panoptic visual recognition and general relation comprehension through region-level annotations and image-text pairs.
Llava recaptioned COCO2014 ValSet. This dataset contains 30,000 images from the COCO 2014 validation set, each paired with its original caption and a new caption generated by the LLaVA vision-language model. It was created by UCSC-VLAA for evaluating text-to-image generation models, as detailed in the associated research paper.
A reformatted version of the GSM8K dataset of grade school math word problems. It presents the original 'socratic' version, where a model reflects by asking sub-questions, as a multi-turn conversation where the sub-questions are posed by a user. The dataset was created by author 'euclaise' and was last updated on Hugging Face in July 2024.
ImageNet-R(endition) contains images of art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video game renditions. The dataset is hosted on Hugging Face by user 'axiong' and was last updated on June 19, 2024. It is constructed from a source file provided by an official implementation to facilitate the evaluation of various pretraining models.
A collection of procurement notices published by the European Union, organized into Parquet files divided by year. The dataset is hosted on Hugging Face by author PleIAs and was last updated on 2024-06-17. Each file contains detailed information about public tenders from the TED platform, the EU's official public procurement journal.
Cnndetection is a dataset uploaded to Hugging Face by user sywang on July 26, 2024. The dataset's title suggests a focus on detecting content generated by Convolutional Neural Networks, likely containing images for classification tasks. Specific details on size, columns, and license are not provided in the available metadata.
VinAIResearch published this scene text recognition collection in 2021 for the CVPR conference. It provides images and dictionary-guided labels for text detection and recognition tasks, specifically targeting the VinText benchmark.
Information on the organizational structure of the communal non-profit medical enterprise 'Center for Primary Medicine - Sanitary Care Ü2' in Kremenchuk, Ukraine. The dataset is defined in accordance with the organizational and administrative document 'Structure and staffing'. It was published on the States site of Ukraine and last updated on June 21, 2024.
TextOCR-GPT4o is Meta's TextOCR dataset captioned with emphasis on OCR using GPT-4o. The dataset is intended for generating benchmarks to compare Vision-Language Models to GPT-4o. CaptionEmporium published it on Hugging Face in June 2024.
61,467 prompt-image pairs collected from the Lexica platform for research on text-to-image generation models. All prompts were curated by real users and the corresponding images were generated by Stable Diffusion. The dataset was shared in a USENIX'24 paper and uploaded to Hugging Face by vera365 on May 16, 2024.