Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
EPIC-KITCHENS-100 is a 100-hour extension of the 2018 EPIC-KITCHENS dataset, containing first-person, non-scripted audio-visual recordings of daily kitchen activities. The dataset was created using a 'Pause-and-Talk' narration interface for annotation.
High-resolution UAV images categorized for the detection of power line assets across multiple size scales. The data supports the development of computer vision models for automated electrical infrastructure monitoring and aerial inspection.
Encompassing a subset of 1,000 Smithsonian butterfly images processed for training a GAN model. It includes a 'sim_score' feature generated by a CLIP model based on descriptive text prompts. The data was curated by 'huggan' and last updated in April 2022.
Featuring images of document forms. It is derived from the CORD dataset and was uploaded by nehruperumalla. The specific number of rows, columns, and file formats is not provided.
A Sentinel-2 satellite image mosaic of the Yanomami Indigenous Territory in Brazil with 10-meter spatial resolution. The mosaic was prepared by INPE's Health Information Investigation Laboratory (LiSS) and ICIT-FIOCRUZ to support a health situation database. It uses a false color composition based on MSI bands 11, 8A, and 4, and encompasses a 6-month temporal composition from April 2019 to September 2022, generated from over 15000 scenes using a Least Cloud Cover First (LCF) best pixel selection approach.
Rice Image Dataset 2 contains images of rice plants for computer vision tasks. The dataset was created by author 'nateraw' and was last updated on Hugging Face in June 2022. Platform tags indicate it contains at least 10,000 to 100,000 images.
A collection of over 100,000 examples of images for training models to correct VQGAN reconstruction artifacts. Each example includes a 512px source image, a 256px version, and a VQGAN-reconstructed version from the 256px image. It was created by johnowhitaker and last updated in April 2022.
Documenting sensor data and pose estimations from an Aubo i5 dual-arm collaborative robot and an Intel RealSense D435 camera. It focuses on 3D object pose estimation using the Deep Object Pose Estimation (DOPE) method within a ROS-compatible structure.
Encompassing image captions from the MSCOCO collection, stored in Parquet format. It was authored by cat-state and last updated in May 2022. The specific row count, column structure, and image data are not detailed in the provided input.
This repository provides scripts for converting computer vision annotation formats between COCO, YOLO, VOC, and XML standards. Developed by MrAsimZahid and last updated in June 2022, it serves as a utility for deep learning dataset preparation.
A subset of the RVL-CDIP collection containing only invoice images. The full RVL-CDIP dataset consists of 400,000 grayscale images across 16 classes, with 25,000 images per class. It includes 320,000 training images, 40,000 validation images, and 40,000 test images.
Encompassing 25,000 macroscopic images of trees from the Peruvian Amazon, representing 25 distinct species. The images were collected and validated by a botanical team using mobile device cameras and digital microscopes, with each image having a resolution of 480x640 pixels and three color channels.
The Flickr-Faces-HQ (FFHQ) dataset contains 70,000 high-quality PNG images of human faces. It was originally created by NVIDIA researchers as a benchmark for generative adversarial networks (GANs).
YT-Temporal-180M is a dataset of 6 million YouTube videos, from which 180 million frames have been extracted. It was created by HuggingFaceM4 and covers diverse topics.
Aggregating 100,000 examples for training models to correct artifacts in VQGAN reconstructions. Each example includes a 512px original image, a 256px downscaled version, and a VQGAN-reconstructed version from the 256px image. It was created by johnowhitaker and last updated in April 2022.
Featuring images from the Mutant Ape Yacht Club NFT collection for unconditional generation tasks. It was created by huggingnft and last updated in April 2022. The dataset is tagged for image modality and generative adversarial network applications.
Giving access to tokenized training and validation data for the BioCreative II gene mention task. It includes sequences of tokens, case-folded tokens, part-of-speech tags, and BIO sequence labels.
Aggregating sample images generated by the TADNE model. It includes prediction results from an anime-face-detector (YOLOv3 + HRNetV2) and deepdanbooru tag predictions, including intermediate 4096-dimensional feature vectors.
A collection of images of art paintings intended for few-shot image synthesis research. It is associated with a 2021 arXiv paper on faster and stabilized GAN training for high-fidelity few-shot image synthesis. The dataset size and specific column structure are not provided in the input.
Aggregating images of Pokemon for few-shot image synthesis research, associated with a 2021 arXiv paper on GAN training. The dataset is categorized as size 'n1 K', indicating it contains over 1,000 images. It is hosted by the 'huggan' organization and was last updated in April 2022.