Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,633 datasets
Over 6,900 images curated for image-to-text recognition tasks, particularly for scanned documents and OCR. The dataset, created by prithivMLmods, is structured in an ImageFolder format and was last updated on September 9, 2025. It is intended for training models on document parsing, PDF image understanding, and layout or text extraction.
A dataset of Turkish words, likely intended for optical character recognition tasks. It was published on the Hugging Face platform by the author 'orkungedik' and was last updated on November 8, 2025. The specific content, size, and structure require verification after download.
Non-profit organizations with addresses in Cambridge, Massachusetts, which have IRS recognition of tax-exempt status. The data includes NTEE classification codes, financial attributes like Assets and Income, and geospatial coordinates. It was last updated on 2025-08-27 by data.cambridgema.gov.
The Register of Geographic Codes (RGC) is the definitive list of UK statistical geographies maintained by the Office for National Statistics (ONS). This version from December 2020 includes late updates for Leeds wards, parishes, and unparished areas. ONS coordinates the issuance of new codes and maintains relationships between active and archived code ranges on behalf of the Government Statistical Service.
Built from annotated images representing six Brazilian Jiu-Jitsu positions: guard, mount, side control, turtle, takedown, and standing. Each image includes category labels, bounding boxes, and keypoints formatted according to the COCO standard across training, validation, and test splits.
A training dataset for the SFE Competition 2025, which appears to be a computer vision challenge. The dataset was uploaded by the author 'InternScience' and last updated on October 10, 2025. The description indicates the data is stored in a zip file bundle named 'images_bundle.zip'.
A balanced collection of AI-generated and real human-captured images for classification tasks. The dataset was created by Parveshiiii and was last updated on September 25, 2025. It contains two classes: '0' for AI-generated images and '1' for real images.
Assembled from protected microdata files from Statistics Netherlands (CBS), processed to prevent the identification of individuals and households. The files are available for free download from the EASY archive, with new versions released regularly. The specific number of rows, columns, and file size is unknown.
A dataset harvested on 2025-10-14 by Begoña Fernández-Pintor via e-cienciaDatos. It contains the volatile composition of seven edible flower species determined by HS-SPME/GC-MS analysis. The data includes information on potential bioactive effects and the number of volatile organic metabolites found.
Replication data for the paper 'Outsourcing of research and development and efficiency: a DEA non-parametric analysis of the contract research organisations industry' by Díaz and Sanchez-Robles (2023). The dataset was harvested from e-cienciaDatos and last updated on October 14, 2025.
Data from the OPTIMSOIL project measuring the Fraction of Photosynthetically Active Radiation intercepted by the canopy (fIPAR), incident PAR, transmitted PAR, intercepted PAR per subperiod, and Radiation Use Efficiency (RUE). The dataset was authored by Daniel Palmero Llamas and last updated on October 14, -2025.
A photograph depicting women producing argan oil using traditional methods. This image is part of the 'Enhanced Tashelhiyt Dictionary: Argan' collection, contributed by author 'Chrumps' and hosted by the DANS Data Station Social Sciences and Humanities Collection. The record was last updated on October 24, 2025.
A photograph of an argan tree, part of the 'Enhanced Tashelhiyt Dictionary: Argan' project. The image was contributed by author Luc Viatour and is hosted by the DANS Data Station Social Sciences and Humanities Collection. It was last updated on October 24, 2025.
Workshop materials from the 'FAIR: Facts and Implementations' event organized by ePLAN and hosted by the Netherlands eScience Center. The materials were authored by P.J.C. Aerts and are part of the DANS Data Station Social Sciences and Humanities Collection. The dataset was last updated on October 24, 2025.
Research data underpinning a study on the unintended side-effects of commercializing smallholder production in Uganda. The dataset was authored by P Ntakyo and is hosted by the DANS Data Station Social Sciences and Humanities Collection. It was last updated on October 24, 2025.
A dataset of high-quality images with precise annotations for object detection tasks. It contains images for 33 different automotive brands, focusing on brand recognition in contexts like vehicles and logos. The dataset was created by haydarkadioglu and was last updated on 2025-09-24.
Imagenet21K Recaption is a dataset of approximately 13 million images across about 19 thousand classes, with labels provided as strings. The dataset was created by author gmongaras and was last updated on the Hugging Face platform on 2025-09-17. Images are stored in PNG format and can be decoded using the Python PIL library as shown in the description.
UniqueData's Grocery Store Receipts Dataset is a collection of photos from various grocery store receipts, designed for Optical Character Recognition tasks. Each image includes bounding box annotations for text segments categorized into four classes: item, store, date_time, and total. The dataset was last updated on October 1, 2025.
A curated subset of 25,000 images from a larger collection of over 1 million soccer images. The dataset was compiled by Adit-jain from multiple open-source sources, including the SoccerNet dataset, video game renders, and match footage, and was last updated on August 28, 2025.
Animepics is a pure image dataset in WebDataset format designed for pre-training anime-style models. The dataset is designed to be continuously updated, leveraging the features of WebDataset. It was authored by zenless-fab and last updated on October 2, 2025.