Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
Identification and contact details for legal entities, individual entrepreneurs, and public associations in Ukraine, as recorded during state registration. The data is maintained by the States site of Ukraine and was last updated on March 8, 2025. The information is defined by the Law of Ukraine On State Registration of Legal Entities, Individual Entrepreneurs and Public Associations.
2520 handwritten text images were curated from native Marathi speakers to ensure variety in handwriting styles and character variations. The dataset supports Optical Character Recognition (OCR) systems, handwriting analysis, and language research. It was created by Process-Venue and last updated in February 2025.
The dataset describes the organizational structure of the Department of Municipal Services and Regulatory Policy of the Kamyanka City Council. It originates from the States site of Ukraine and was last updated on March 13, 2025. The available file formats include ODS, XLSX, XML, and CSV.
An open-source dataset of 25,000 annotated images for object detection in agricultural applications. The dataset, created by devshaheen, contains images of 100 different crops and plants and was last updated on February 25, 2025.
30,000 high-resolution face images selected from the CelebA dataset, following the CelebA-HQ selection process. Each image includes a corresponding segmentation mask for facial attributes. The dataset, created by author cpuimage, was last updated on March 1, 2025.
Hundreds of thousands of user-submitted prompts captured during the Mosscap prompt injection game at DEF CON 31. The collection includes raw text inputs ranging from successful adversarial attacks to general conversational queries directed at the AI character across multiple challenge levels.
A collection of prompt injection attempts filtered from millions of user interactions with the Gandalf security game during July 2023. The data specifically isolates adversarial inputs designed to bypass system instructions, distinguishing them from general user queries through automated text analysis.
ZHOUYI1023 maintains this curated repository of radar perception resources, updated through April 2025. It organizes links to external datasets, detection algorithms, and sensor fusion techniques specifically for autonomous driving applications.
CVRPDataset created a collection of 2,303 field images of rice plants, captured from 231 landraces and 50 modern cultivars. Each image is paired with an annotated mask. The dataset also includes 123 indoor images focused on individual panicles and 3D target reconstruction files, with the last update recorded on 2025-02-28.
Comprising 13,513 high-quality images from five significant earthquakes, labeled for infrastructure damage assessment. It uses a 4-class classification system aligned with established damage scales like HAZUS and EMS-98, with labeling guidelines that improved annotator agreement from 39.7% to 10.4% disagreement.
The dataset from the States site of Ukraine contains identifiers, names, descriptions, and operational details for fairs. It includes information on rental costs, addresses, duration, work schedules, organizers, and their contacts, as well as contracts concluded with organizers. The data was last updated on March 8, -2025.
Information about the organizational structure of the Qualification and Disciplinary Commission of the Bar of Ivano-Frankivsk region. The dataset is provided by the States site of Ukraine and was last updated on March 7, 2025. It is available for download in the EXCEL XLSX file format.
Handwritten homework samples from pupils across Chinese, Mathematics, and English subjects. The dataset, authored by XinyueZhou, was last updated on February 17, 2025. It is structured to allow alignment of work across different subjects for each student.
Schematic to HSPICE Netlist Dataset is designed for converting schematic images into HSPICE netlists. The dataset, created by hanky2397, includes images and component annotations in YOLO format for object detection and net extraction tasks. It was last updated on February 20, 2025.
Flickr-Faces-HQ (FFHQ) images downsampled to 32x32 resolution. The original full-resolution dataset was created by NVIDIA Research and published in the 2019 CVPR conference. This downsampled version was uploaded to Hugging Face by user 'leellodadi' on March 13, 2025.
Deepfake vs Real is a dataset for image classification, distinguishing between deepfake and real images. It includes a diverse collection of high-quality deepfake images and aims to support the development of more robust detection models. The dataset was created by prithivMLmods and was last updated on February 19, 2025.
The U.S. Environmental Protection Agency's Sustainable Materials Management (SMM) Electronics Challenge tracks electronics manufacturers, brand owners, and retailers committed to sending 100% of collected used electronics to certified refurbishers and recyclers. The program publicly recognizes participants as registrants, new participants, or active participants and offers tier and champion awards. Specific data points include participant status and award categories.
SSCBench is a large-scale benchmark for 3D semantic scene completion in autonomous driving scenarios, presented at IROS 2024 by the ai4ce research group. The dataset provides 3D occupancy grid maps and semantic labels to facilitate 3D scene understanding from 2D or 3D sensor inputs. It was released in 2024 to address the gap in standardized evaluation for semantic scene completion tasks.
VisionArena-Battle contains 200,000 single and multi-turn chats between users and 45 different Vision Language Models (VLMs), collected via the Chatbot Arena platform. The dataset includes approximately 43,000 unique images and conversations across 138 languages, with question categories such as captioning, OCR, and creative writing. It was created by lmarena-ai and last updated on February 4, 2025.
A dataset of ImageNet 1K images recaptioned using the Moondream2 vision-language model. The author 'g-ronimo' processed the data on March 9, 2025, generating short, precise captions based on the original ImageNet class names and existing captions. The dataset contains only image IDs and the newly generated text captions.