Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
The organizational structure of the Department of Economics and Development within Cherkasy City Council. The dataset is provided by the States site of Ukraine and was last updated on June 20, 2025. It likely contains information about the department's internal hierarchy and personnel roles.
400,000 synthetic graphic design samples, each annotated with multiple conditions for generation tasks. The dataset includes primary and secondary visual elements with spatial and semantic annotations, plus textual elements with layout information. It was created by HuiZhang0812 and was last updated in June 2025.
Over 2 billion public messages were collected from 3,167 distinct public Discord servers, accompanying a paper submitted to ICWSM 2025. The dataset is organized into individual JSON files per server, with an overview file providing server metadata. It was uploaded by author fvdfs41 and last updated on June 9, 2025.
CAGUI is a benchmark dataset for evaluating GUI agent models on Chinese Android applications. It focuses on two capabilities: grounding GUI components to semantics and planning multi-step actions to complete user goals. The dataset was created by openbmb and last updated in June 2025.
This repository provides label mapping files that convert integer IDs to human-readable class names for multiple datasets. It includes mappings for ImageNet-1k, ImageNet-22k (with 21,843 classes), COCO detection 2017, and others like VQAv2 and Kinetics-700. The files are maintained by HuggingFace and were last updated in June 2025.
Finevd is a dataset hosted on HuggingFace by IntMeGroup, last updated on July 15, 2025. Access requires contacting the author via email with details about the requester's name, organization, and intended use. The dataset's specific content, size, and structure are not publicly documented.
GOES-16 ABI satellite image dataset with multi-spectral imagery and corresponding labels. The dataset contains training and test splits at two different resolutions, 128x128 and 256x256. Data was provided by NOAA and NESDIS and uploaded by Silicon23 in July 2025.
Over 10,000 individuals in New York are impacted by tick-borne diseases yearly. The New York State Department of Health (NYSDOH) monitors pathogens through active tick surveillance, which began statewide in 2008. This dataset records nymph tick density and infection prevalence across NY regions from May to September.
Aesthetic-Train-V2 is a high-quality training set for ultra-high-resolution image generation. The dataset was introduced by author zhang0jhon in the CVPR 2025 paper 'Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models'. The dataset was last updated on the Hugging Face platform on June 4, 2025.
The dataset contains information about the organizational structure of the UPSZN Ladyzhyn City Council and the number of staff units. It was published on the States site of Ukraine and last updated on June 18, 2025. The data is available in Excel, XML, and Word formats.
2 images sourced from the IAM Handwriting Database and the SRIOE dataset. These files serve as functional test fixtures for validating TrOCR vision-encoder-decoder models within the HuggingFace Transformers library.
OrgAccess is a synthetic benchmark dataset created by respai-lab to evaluate the ability of Large Language Models to operate within organizational hierarchies and role-based access control policies. The dataset was last updated on June 12, 2025. It addresses a critical challenge in ensuring LLMs can reliably function as unified agents within structured organizations.
AutomotiveUI-Bench-4K is a validation benchmark dataset containing 998 images and 4,208 annotations focused on user interaction with in-vehicle infotainment systems. The dataset was created by sparks-solutions and covers 15 automotive brands for model years 2018-2025. It was last updated on Hugging Face on May 14, -2025.
MicroG-HAR-train-ready is a dataset converted from the MicroG-4M dataset for human activity recognition. The dataset is formatted to be used directly as input data directories for the PySlowFast_for_HAR framework without additional preprocessing. It was authored by lei-qi-233 and last updated on June 9, 2025.
179 equirectangular RGB images with corresponding depth, surface normals, XYZ, and HHA images. The dataset, created by COLE-Ricoh, provides instance-level semantic and room layout annotations for 4 unique scenes. It was last updated on Hugging Face in June 2025.
AfriMed-QA v2 is a novel multispecialty medical question-answering dataset for Africa. The dataset was created by a collaboration including Intron Health, SisonkeBiotik, BioRAMP, Georgia Institute of Technology, MasakhaneNLP, and Google Research, with funding from Google Research, the Bill & Melinda Gates Foundation, and PATH. It was last updated on 2025-06-17.
COCO 2017 Mirror is a hosted copy of the original COCO 2017 dataset files, provided by author pcuenq for convenient access. The dataset is tagged for tasks including Image, Computer Vision, Object Detection, and Image Captioning. It was last updated on July 4, 2025.
Japanese Synthetic OCR 150K is a dataset of synthetic images for optical character recognition tasks. The dataset was created by the author 'deepcopy' and was last updated on July 7, 2025. Its source is listed as a Kaggle dataset containing Japanese font images.
A curated collection of text data in English, French, German, Spanish, and Italian, culled from sources including web data, video subtitles, academic papers, digital books, newspapers, and magazines. The dataset was used to pretrain the Lucie-7B foundation LLM and also contains samples of diverse programming languages. It was authored by OpenLLM-France and last updated on the Hugging Face platform in May 2025.
CAFOSat is a remote sensing dataset for identifying Concentrated Animal Feeding Operations (CAFOs) across various U.S. states. It includes high-resolution image patches, infrastructure annotations, bounding boxes, and experimental train-test splits for multiple configurations. The dataset was created by oishee3003 and was last updated on June 6, 2025.