Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
ICVL is a hyperspectral image dataset collected for the research 'Sparse Recovery of Hyperspectral Signal from Natural RGB Images'. It contains at least 200 images captured with a Specim PS Kappa DX4 camera and a rotary stage, with a spatial resolution of 1392 x 1300 pixels. The dataset was last updated on October 1, 2024, by the author danaroth.
19th century texts published in Russia using pre-reform orthography. The dataset contains source images paired with human-readable extracted texts, designed to train and evaluate optical character recognition systems for Russian texts published before the 1917 orthographic reform. It was created by user 'nevmenandr' and last updated on Hugging Face in October 2024.
A dataset from the paper "Abductive Ego-View Accident Video Understanding for Safe Driving Perception" by JeffreyChou, uploaded to Hugging Face on 2024-10-09. The data is chunked and compressed, requiring merging and extraction after download, as demonstrated with the DADA2000 example. The dataset is intended for research in understanding driving accidents from a first-person perspective.
A dataset for the Pokemon Trading Card Game (TCG) containing information about cards, decks, and sets. It was created by the author 'tooni' and last updated on October 20, 2024. The dataset is designed for training machine learning models.
Bathymetry for Narragansett Bay, Rhode Island was derived from fifteen hydrographic surveys containing 165,184 individual soundings. The data provides a Digital Elevation Model with 30-meter grid spacing, covering the estuary with thirteen 7.5-minute DEMs. Surveys used date from 1943 to 1957, with soundings ranging from 1.2 meters to -58.8 meters relative to mean low water.
116k text samples were created by Stanley Sebastian of Replete-AI to explore the true capabilities and limits of large language models. The dataset is positioned as a tool to move beyond fine-tuning models for assistant roles and instead examine them as collections of memorized semantic patterns. It was last updated on Hugging Face on October 20, 2024.
FashionFail is a dataset of 2,495 high-resolution (2400x2400 pixels) images of e-commerce products, proposed in the paper "FashionFail: Addressing Failure Cases in Fashion Object Detection and Segmentation". The dataset is split into 1,344 training images, 150 validation images, and 1,001 test images. It was authored by rizavelioglu and last updated on Hugging Face in October 2024.
MMBench-Video is a quantitative benchmark designed to rigorously evaluate Large Vision-Language Models' proficiency in video understanding. It incorporates approximately 600 web videos and was created by opencompass. The dataset was last updated on October 9, 2024.
Labeled Faces in the Wild is a database for studying face recognition in unconstrained environments. It was created by Gary B. Huang, Manu Ramesh, Tamara Berg, and Erik Learned-Miller at the University of Massachusetts, Amherst. The technical report was published in October 2007.
FreedomIntelligence created a multilingual dataset for medical large language models, covering 12 major and 38 minor languages. The dataset was last updated on October 18, 2024. Its primary purpose is to democratize medical LLMs across a wide range of languages.
E.T. Instruct 164K is a large-scale instruction-tuning dataset for fine-grained event-level and time-sensitive video understanding. It contains 101,000 videos across diverse domains and 9 event-level understanding tasks with instruction-response pairs, with an average video length of around 146 seconds. The dataset was created by PolyU-ChenLab and last updated on Hugging Face in September 2024.
HuMMan is a multi-modal 4D human dataset presented at ECVV 2022. The dataset is hosted by author caizhongang on HuggingFace and was updated on October 7, 2024. It includes subsets for motion generation (HuMMan-MoGen), 3D vision (HuMMan-Point), and depth maps (HuMMan-Recon).
20,005 person sequences in the GTA-Human dataset and 35,352 sequences in GTA-Human II, released in 2022 and 2024 respectively, provide synthetic data for 3D human recovery. The datasets were created by caizhongang and are hosted on HuggingFace. They are associated with research published in TPAMI 2024.
AI-TOD provides aerial imagery for tiny object detection, released by jwwangchn and updated in November 2024. It focuses on identifying objects with extremely small pixel footprints in high-altitude captures.
The 'Welsh Texts' dataset contains a variety of printed and handwritten material from Welsh sources, mostly in the Welsh language. The dataset includes works such as 'Drych y Prif Oesoedd' by Theophilus Evans, a book on the early history of Wales published in 1716, and 'Enwogion Cymreig' by Thomas Morgan, cataloging prominent figures. It is hosted by OpenAI on Hugging Face, with authorization from The National Library of Wales and The Welsh Government.
SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions). It was created by author meganwei and last updated on Hugging Face on October 2, 2024. The dataset includes 7 distinct concepts, each with its own configuration.
A repackaged version of the ILSVRC/ImageNet-1k dataset, transformed into a uniform 64x64 pixel resolution. The images were center-cropped to square and resampled using the Lanczos algorithm, with higher-resolution versions also available. This version was created by benjamin-paine and last updated on September 15, 2024.
ImageNet 1K 32X32 is a repackaged version of the ILSVRC/ImageNet-1k dataset. The images have been center-cropped to square and resampled to a 32x32 pixel resolution using Lanczos interpolation. The dataset was created by benjamin-paine and last updated on Hugging Face in September 2024.
benjamin-paine repackaged the ImageNet 1K dataset into a 128x128 pixel format. Images were center-cropped and resampled to this uniform size using Lanczos interpolation. The repack was last updated on September 15, 2024.
SkyScenes is a collection of synthetic, densely annotated images designed for aerial scene understanding across a diverse set of environmental conditions. The dataset provides pixel-level semantic labels to address the challenges of data scarcity and annotation costs inherent in real-world aerial photography.