Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,713 datasets
Synthetic_Luganda_VITS_22.5k provides 22,500 audio samples for a low-resource language. The dataset was created by mekaneeky and uploaded to Hugging Face in October 2023. It likely contains synthetic speech audio generated for Luganda language applications.
PubTables-1M is a dataset for evaluating object detection and image-to-text models on table extraction from unstructured documents. It was introduced by Smock et al. and converted to Hugging Face format with an Optimized Table Structure Language (OTSL) by the docling-project. The dataset page was last updated on 2023-08-31.
A conversion of the original SynthTabNet dataset into the OTSL format, as presented in the paper 'Optimized Table Tokenization for Table Structure Recognition'. The dataset includes original annotations and new additions, organized into 4 parts totaling 600,000 tables with varied appearances. It was created by docling-project and last updated on Hugging Face in August 2023.
A dataset of scanned invoice and receipt images intended for optical character recognition (OCR) tasks. It was uploaded to Hugging Face by the user 'mychen76' and last updated on September 22, 2023. The specific number of images, their geographic origin, and collection method are not detailed in the available metadata.
An evaluation dataset for optical character recognition (OCR) systems, focusing on receipts. The dataset was created by author mychen76 and was last updated on Hugging Face in September 2023. Full details are referenced in a related training dataset.
Arabic OCR data and tools updated in October 2023 by HusseinYoussef focus on character segmentation and neural network recognition for typed text. The repository provides image-based script samples for transformation into machine-encoded text using OpenCV and scikit-learn.
30,000 chemical structure images and corresponding graph-based labels curated from United States Patent and Trademark Office (USPTO) documents. The dataset serves as a benchmark for Optical Chemical Structure Recognition (OCSR) and was introduced alongside the MolGrapher architecture to address data redundancy in existing chemical vision tasks.
An Indonesian-language instruction-tuning dataset derived from the FreedomIntelligence/alpaca-gpt4-indonesian model. The dataset was reformatted into 'input' and 'output' pairs for supervised fine-tuning. It was created by user Ichsan2895 and last updated on Hugging Face in August 2023.
10,738 unique music clips paired with 32,214 images across 13 distinct emotional categories derived from subjective experience studies. Each music entry is linked to three images sharing the same emotional label to support cross-modal sentiment analysis and alignment tasks.
32,203 images containing 393,703 labeled face instances across various scales, poses, and occlusion levels. These annotations are provided in the Pascal VOC XML format, facilitating integration with object detection frameworks like Faster R-CNN or SSD.
Aggregate data from the Connecticut Department of Children and Families on abuse and neglect reports accepted for response. The dataset tracks reports handled through both mandated Child Protective Services investigations and the voluntary Family Assessment Response process implemented in April 2012. It includes information on allegations, substantiations, and service provision by town and state fiscal year.
IDEA-Research released Human-Art in 2023 to provide a multi-scenario human-centric collection bridging natural and artificial visual domains. The data supports human pose estimation and image generation across diverse artistic styles and real-world settings.
MaryamBoneh released this vehicle detection resource in September 2023, focusing on object detection using the YOLOv5 algorithm. The project provides image data and annotations specifically designed for fine-tuning deep learning models to perform car-counting and vehicle-tracking tasks.
Imagenet_1k_resized_256 is a derivative of the ImageNet dataset where the smaller side of each image has been resized to 256 pixels. The dataset was created by user 'evanarlian' and was last updated on Hugging Face on 2023-08-01. This resized version is intended to speed up download times and reduce storage requirements compared to the original dataset.
Information on the organizational structure of the Department of Environmental Protection and Energy of Sumy Regional State Administration. The dataset was published on the States site of Ukraine and last updated on 2023-08-15 06:09:16.541118. It is available for download in EXCEL XLS and EXCEL XLSX formats.
1,000,000+ annotations for vehicle detection and orientation classification. This dataset enables the training of standard object detection networks to predict both bounding boxes and vehicle headings.
COCO-AB is an extension of the COCO 2014 training set, enriched with additional annotation byproducts. The dataset contains 82,765 reannotated images from the original set. It was authored by coallaoh and last updated on Hugging Face in July 2023.
123,287 images across training and validation sets featuring object detection, segmentation, and captioning annotations. The collection includes 118,287 training images and 5,000 validation images with labels for 80 distinct object categories.
1 PyTorch implementation of the AlexNet architecture featuring 5 convolutional and 3 fully connected layers. The code provides the structural definition for the 2012 ImageNet-winning model within the torch.nn framework.
An object detection dataset focused on elements in mobile user interface designs. The dataset contains images with bounding boxes and class labels for objects including text, images, and groups. It was created by author mrtoy and last updated on Hugging Face in July 2023.