Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Scraped from the official International Standardization Organization (ISO) website, this dataset provides metadata on international standards as of November 2024. It was compiled by user sbjorkholt to support research in internationalization, social science, and political science.
A collection of 9.3 million document images processed through Optical Character Recognition (OCR) using the docTR library. The dataset is derived from the AMF-PDF dataset, part of the Finance Commons collection, and was created by lightonai. The dataset card was last updated on September 23, 2024.
A dataset for evaluating Optical Character Recognition models on Thai language text. It contains images and corresponding textual data derived from various open-source websites. The dataset was created by openthaigpt and last updated in September 2024.
An enriched version of the ImageNet-1K dataset, last updated on 2024-09 16. It includes image captions, bounding boxes, and label issues, extending the dataset's utility beyond classification. The dataset was created by visual-layer and is hosted on Hugging Face.
AMEX provides multi-annotated Android GUI data for training mobile agents, released in 2024 by Yuxiang007. The dataset includes screenshots paired with instruction-action chains and fine-grained element metadata such as bounding boxes and functional descriptions.
Deeva provides analytics and visualization metadata for object detection datasets, developed by vbyan and last updated in November 2024. It generates statistical insights and interactive Plotly visualizations for machine learning data using a Streamlit-based interface.
VGGFace2-HQ is a high-resolution version of the VGGFace2 dataset created for academic face editing. The dataset was generated using GFPGAN for image restoration and insightface for data preprocessing, including crop and align operations. It is hosted on Hugging Face by RichardErkhov and was last updated on September 25, 2024.
A dataset supporting the paper "Diffusion Models for Monocular Depth Estimation: Overcoming Challenging Conditions" (ECCV 2024). It is organized into three main categories: autonomous driving scenes with challenging images, a Transparent and Mirrored (ToM) objects dataset, and trained model weights. The dataset was uploaded by fabiotosi92 and last updated on September 28, 2024.
ImageNet 1K is a repackaged version of the ILSVRC 2012 dataset containing images organized according to the WordNet hierarchy. The images have been center-cropped to square and resampled to a uniform 256x256 pixel resolution using Lanczos resampling. This specific repack was created by benjamin-paine and uploaded to Hugging Face in September 2024.
An enriched version of the COCO 2014 dataset. It includes label issues to help curate a cleaner dataset. The dataset was uploaded by visual-layer and last updated on September 16, 2024.
A research dataset investigates physiological and genetic factors influencing reproductive success in Weddell seal females. The study, conducted by AMD_USAPDC and published in 2024, employs an organismal energetics and genome-to-phenome approach to compare high- and low-quality individuals. It integrates repeated physiological sampling with whole-genome sequencing data.
Filled with 50,000 images, with 50 images for each of the 1,000 ImageNet classes. It was constructed by querying Google Images for sketches of each class name and manually cleaning the results.
Imagenette is a subset of 10 easily classified classes from the ImageNet dataset, including tench, English springer, and cassette player. It is designed for researchers and practitioners to quickly test and share ideas, and is available in three image size variants: full size, 320 px, and 160 px.
KennethWussmann released Caption.Now in November 2024 as an offline-first Progressive Web App (PWA) for manual image captioning. The tool facilitates the creation of text-image pairs and classification labels required for training AI models.
RAID is a benchmark dataset for evaluating machine-generated text detectors, containing over 10 million documents. It spans 11 large language models, 11 genres, 4 decoding strategies, and 12 adversarial attacks. The dataset was created by liamdugan and updated in September 2024.
A dataset of handwritten text images for optical character recognition tasks. It was published on HuggingFace by AadityaJain and was last updated on October 31, 2024. The specific content, scale, and collection method are not detailed in the available metadata.
32,043 brightfield microscopy images sliced from 4K resolution sources, featuring cell-level bounding box annotations. The collection is partitioned into 27,161 training, 3,302 validation, and 1,580 test examples for biological imaging tasks.
Information about fairs, their organizers, and related contracts published by the States site of Ukraine. The data likely contains details such as event terms, locations, seat costs, and contractual agreements. It was last updated on October 2, 2024, and is available in Excel XLSX format.
Advertising images and human feedback annotations across commercial categories comprise this ECCV 2024 dataset. It focuses on human-perceived reliability to improve the visual fidelity and brand alignment of synthetic marketing content.
Filled with 50,000 images, with 50 images for each of the 1,000 ImageNet classes. It was constructed by querying Google Images for sketches of each class name and manually cleaning the results.