Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
The RVL-CDIP dataset contains 400,000 grayscale images of documents, evenly distributed across 16 classes with 25,000 images per class. It is split into 320,000 training, 40,000 validation, and 40,000 test images. The dataset was created by aharley for document information processing tasks.
6,745 images across train, validation, and test splits for classifying 20 different snack foods. The images were sourced from the Google Open Images dataset release 2017_11 and accompany a machine learning tutorial book. It includes categories such as apple, banana, cake, candy, and carrot.
An object detection dataset of powerline infrastructure images with annotated bounding boxes for components and faults. The dataset was created by author docmhvr and was last updated on August 30, 2024. It includes mosaic augmentation for training and evaluating models.
National Oceanic and Atmospheric Administration provides multibeam bathymetry data collected aboard the Northwestern vessel from 15-May-23 to 19-May-23. This dataset is part of the larger Multibeam Bathymetry Database (MBBDB). The survey covers a route from Traverse City to Northport, Michigan.
Magpie-Align's filtered 300,000 instruction-response pairs aim to democratize AI by providing open alignment data, as described in their technical report from June 2024. The dataset, last updated on August 28, 2024, is intended to address the high cost and limited scope of human-generated prompts for aligning large language models. It serves as an open alternative to the private alignment data used for models like Llama-3-Instruct.
The International Skin Imaging Collaboration published 25,331 dermoscopic images across nine diagnostic categories. The dataset, hosted by MKZuziak on Hugging Face, serves as a challenging benchmark for medical image classification. The resized version was last updated on August 6, 2024.
City of Detroit zoning data from 2010 details the location of zoning codes, with parcels sharing the same code dissolved together but separated by street boundaries. The dataset includes an attribute table with zoning code descriptions. Data is provided by the City of Ferndale, Michigan, and was last updated in September 2024.
Imagenette is a subset of 10 easily classified classes from the ImageNet dataset, originally prepared by Jeremy Howard of FastAI. It is designed for rapid experimentation and comes in three image size variants: full size, 320 px, and 160 px.
9,003 samples of online programmatic ad creatives include ad sizes and extracted creative text. The dataset, created by PeterBrendan, contains various ad dimensions like (300, 250) and (728, 90). It was last updated in August 2024.
Magpie-Align released this dataset on 2024-08-28. It contains filtered instruction data intended for aligning large language models, specifically targeting models like Llama-3-Instruct. The dataset was created to address the lack of open alignment data for such models.
Magpie-Align released this dataset on 2024-08-28. It contains instruction data for aligning large language models, specifically targeting the Llama-3.1-Pro model. The dataset was created to address the lack of open alignment data for models with open weights, as detailed in the associated technical report.
Magpie-Align released the Magpie Qwen2 Pro 200K Chinese dataset to advance the democratization of AI alignment data. The dataset contains 200,000 high-quality instruction-response pairs in Chinese, intended for aligning large language models. It was published on Hugging Face on August 22, 2024.
An augmented version of the Amazon Shopping Queries Dataset, SQID includes image information for over 190,000 products. The dataset was created by crossingminds and last updated on Hugging Face in September 2024. It pairs product search queries from real Amazon users with up to 40 potentially relevant results.
Offering a non-exhaustive collection of images generated by the DALL-E Mega model for unconditional image generation. It was created by TheBirdLegacy and is categorized under the 'unconditional Image Generation' task with a size between 1,000 and 10,000 samples. The dataset was last updated on September 25, 2024.
230,000 total samples for Persian optical character recognition, split into 184,000 training and 46,000 test examples. The dataset was uploaded by user 'ordaktaktak' to Hugging Face and was last updated on September 26, 2024. It likely contains images of Persian text paired with corresponding transcriptions.
RitAreaSciencePark published this dataset on Hugging Face on 2024-10-09. It appears to be a subset of the ImageNet collection, focusing on 100 distinct visual categories. Each class likely contains 100 image samples, suggesting a structured benchmark for computer vision tasks.
1.27 million video-instruction pairs across categories such as video object grounding, tracking, and captioning. The content focuses on fine-grained object-level perception by mapping visual features to spatial-temporal coordinate tokens within video sequences.
3 million instruction-response pairs were synthetically generated to align open-weight large language models. The Magpie-Align team created this dataset to democratize access to high-quality alignment data, releasing it in August 2024. It addresses the gap left by proprietary alignment data from models like Llama-3-Instruct.
PreFLMR M2KR is a benchmark dataset for multimodal knowledge retrieval. The uploaded collection contains all images used in the M2KR benchmark for training and evaluating models. The dataset was uploaded by BByrneLab and last updated on September 18, 2024.
5,000 raw text prompts extracted from Midjourney's Discord server. The collection consists of uncleaned user-generated strings used to trigger generative AI image creation across various styles and subjects.