Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
MetaShift is a dataset of datasets containing 12,868 sets of natural images across 410 classes. It was created by jameszou707 for evaluating machine learning model performance across diverse data distributions.
SVHN contains over 600,000 labeled images of house numbers captured from Google Street View, developed by researchers at Stanford University. It serves as a real-world benchmark for digit recognition that incorporates natural variations in lighting, font, and background noise.
PIPPA Scored is a subset of the PIPPA dataset annotated with GPT-4 scores for 19 personality traits and 5 meta-attributes. The dataset was created by IlyaGusev and last updated on Hugging Face on August 7, 2024. Each attribute includes a textual explanation generated by ChatGPT.
A corpus of legislative proceedings from California, Texas, New York, and one other large U.S. state between 2015 and 2018. The dataset was created by Khosmood, Dekhtyar, Ellwein, and White and presented at Digital Humanities in Washington, DC in August 2024. It is distributed under the CC BY-NC-SA 4.0 license.
A combined and deduplicated version of the COCO-2014 and COCO-2017 datasets for object detection. The dataset features labels in Turkish and is formatted for instruction-tuning with separate prompt and completion columns. It was created by user 'ucsahin' and last updated on August 6, 2024.
A collection of video frames from AI-generated videos and real sources for training synthetic video detectors. The dataset includes frames from seven AI generation algorithms and real video frames from the VideoASID dataset. It was created by author ductai199x and uploaded to Hugging Face in August 2024.
M4-Instruct is a collection of multi-image datasets constructed for training large multimodal models. It was collected in April 2024 and released in June 2024 by lmms-lab. The data is sourced from public datasets and generated via the GPT-4V API.
Containing between 1,000 and 10,000 records of financial documents for OCR tasks, released by user mychen76 in September 2024. It provides paired image and text data for invoices and receipts, formatted as Parquet files for use with the Hugging Face datasets library.
Mireu-Lab provides a processed version of the NSL-KDD dataset, a benchmark for network intrusion detection systems. The data was converted from ARFF to CSV format and stored with float64 data types. The dataset was last updated on HuggingFace in July 2024.
9 million image URLs across training, validation, and test sets annotated with four distinct computer vision labels. The data includes image-level labels, bounding boxes, object segmentation masks, and visual relationships curated by Google LLC.
Loaded with 50 million drawings contributed by players of the game Quick, Draw!. It covers 345 distinct categories of sketched objects.
A dataset likely intended for image captioning tasks, as suggested by its title. It was uploaded to the Hugging Face platform by the author Luna288 and was last updated on September 4, 2024. The specific content, size, and structure of the data are not described in the available metadata.
Official documentation from the World Trade Organization sourced from its Documents Online platform. The dataset includes documents in the three official languages from 1995 onwards and was uploaded by author PleIAs. The dataset card was last updated on July 12, 2024.
UCF101 is an action recognition dataset containing 13,320 realistic videos across 101 action categories, collected from YouTube. The dataset was uploaded to Hugging Face by flwrlabs and was last updated on July 16, 2024. This specific version stores video frames as individual images, with train and test splits based on the original authors' lists.
The Department of the Treasury provides searchable information on organizations' tax-exempt status and filings. It includes data on eligibility for tax-deductible contributions, Form 990 Series Returns, and Automatic Revocation of Exemption lists. The dataset was last updated in August 2024.
A database of funding awards from the Best Starts for Kids initiative in King County. It details financial contracts awarded to community partners for programs supporting child, youth, and family well-being. The data is published by data.kingcounty.gov and was last updated in June 2024.
A dataset for object detection tasks in the fashion domain, published on the Hugging Face platform by author yainage90. The dataset was last updated on August 22, 2024. The specific content, scale, and annotation details require verification after download.
20 object classes are annotated in the Pascal Visual Object Classes (VOC) dataset, a widely used benchmark for computer vision tasks. The dataset is designed for object detection, image classification, semantic segmentation, and action classification. It was uploaded to Hugging Face by user merve on 2024-07-06.
Over 18,000 patent images from more than 7,000 European patent applications filed in 2020. This curated collection supports research in image captioning, abstract reasoning, and automated document processing. The dataset was created by danaaubakirova.
Annual IRS extract of Form 990 financial data for tax-exempt organizations, excluding private foundations. The dataset is produced by the Department of the Treasury for public release. Specific row and column counts are not provided in the input.