Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
Distributed Acoustic Sensing data from the PoroTomo Natural Laboratory includes raw SEG-Y and HDF5 files from horizontal and vertical fiber optic deployments. Data collection occurred in March 2016, with files organized into 30-second and daily intervals. The dataset is hosted by the Department of Energy on AWS via the Open Energy Data Initiative.
PubTables-1M_OTSL-v1.1 is a filtered version of the PubTables-1M dataset, containing tables enriched with header information. The dataset is designed for evaluating object detection models and image-to-text methods. It was created by docling-project and last updated on 2025-02-10.
DAILtech released this collection of public natural image classification datasets on GitHub, with the latest update occurring on April 22, 2025. The repository aggregates resources specifically for computer vision and deep learning tasks involving natural imagery. It serves as a centralized point for accessing various image-level classification benchmarks.
99,000 multimodal records across audio, image, text, and tabular formats were released by lmms-lab in March 2025 for egocentric life-logging research. The dataset is associated with the EgoLife project and the arXiv:2503.03803 publication.
A text dataset derived from organic chemistry PDFs, likely containing extracted chemical terms and nomenclature. It was published by mlfoundations-dev on the Hugging Face platform and last updated on March 28, 2025. The specific content and scale require verification after download.
Ganjoor Chunked ASR Dataset is a text collection likely derived from the Ganjoor Persian poetry repository, published on HuggingFace by the user 'farsi-asr'. The dataset was last updated on March 23, 2025, suggesting it is a recent resource. Its title indicates it is chunked, which may imply the text is segmented for processing by automatic speech recognition models.
Persian poems with their corresponding meters, sourced from ganjoor.net. The dataset was uploaded by author cnababaie and was last updated on March 15, 2025. The specific scale and structure of the data are not detailed in the provided metadata.
A reorganized video dataset derived from the Wild-Heart/Disney-VideoGeneration-Dataset. It was created by svjack and uploaded to Hugging Face on March 14, 2025. The dataset is intended for fine-tuning the Mochi-1 model.
Sphar contains 7,759 spatio-temporally cropped videos categorized into 14 human action classes, curated by AlexanderMelde and updated in April 2025. The collection aggregates footage from multiple sources specifically filtered to simulate surveillance-camera viewpoints.
Auto-ACD is a large-scale, high-quality dataset containing over 1.9 million audio-text pairs. The dataset builds on the audio-visual correspondence found in existing video datasets like VGGSound and AudioSet. It was created by Loie and is hosted on Hugging Face, with a paper published in 2023.
Approximately 380 billion tokens of government text and data aggregated from open data programs. The Open Government dataset is curated by AgentPublic and was last updated on January 31, 2025. It currently features collections from the US, France, European, and international organizations, structured through the Finance Commons and Legal Commons initiatives.
ViOCRVQA is a dataset for Vietnamese Optical Character Recognition and Visual Question Answering. It contains over 28,000 images and 120,000 question-answer pairs focused on text appearing in images. The dataset was created by huyhuy123 and was last updated on Hugging Face in February 2025.
A mirrored dataset of 18 robot manipulation tasks originally hosted by the PerAct project. The data is organized into train, validation, and test splits. It was uploaded to Hugging Face by hqfang on 2025-02-07 to provide easier access than the original Google Drive source.
Annual data for over 300 health indicators organized into 15 health topics, developed by New York State in 2012. The CHIRS provide data for all 62 New York State counties, 11 regions, and the state overall. Data is sourced from health.data.ny.gov and was last updated in December 2024.
37 specific cat and dog breeds are represented in this multi-modal dataset. It contains images with corresponding segmentation masks and bounding boxes, created by the author 'cvdl'. The dataset was last updated on the Hugging Face platform on February 6, 2025.
ByteDance released this time series dataset on Hugging Face in February 2025. The data is organized in the TFB format, which uses a three-column long table structure. The first column is a required date field, with timestamps that can be in string, datetime, or other formats.
Approximately 350 health indicators across 15 topics are tracked for all 62 New York counties and 11 regions. The data includes trend values and three-year averages, updated regularly by health.data.ny.gov since 2012. It consolidates information from the County Health Assessment Indicators (CHAI) for statewide community health monitoring.
A collection of 150,000 synthetic instruction-response pairs for aligning large language models, created by the Magpie-Align project. The dataset was released in 2024, as indicated by the associated technical report, and is hosted on Hugging Face. It aims to provide an open alternative to private alignment data used by models like Llama-3-Instruct.
A directory from the State site of Ukraine provides listings for enterprises, institutions, and organizations under the management of an information manager. The set contains three resources: organizations (legal entities), organizationUnits (structural subdivisions), and posts (officials/employees). It was last updated on January 30, 2025.
LogoDet-3K is a dataset of thousands of images containing brand logotypes and their corresponding bounding boxes. The dataset was created by axonstan and aims to support the training of logotype detection models. It was last updated on the Hugging Face platform on February 3, 2025.