Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
Magpie-Align released this dataset of 250,000 instruction-following examples to democratize AI by providing open alignment data for large language models. The data is designed to improve reasoning capabilities, addressing the gap left by proprietary datasets from models like Llama-3-Instruct. It was last updated on January 27, 2025.
OmniAI OCR Benchmark provides between 1,000 and 10,000 image-text pairs for evaluating multimodal LLM performance, released by getomni-ai in February 2025. It measures the accuracy of text and JSON extraction from images, specifically targeting models like GPT-4o and Gemini 2.0.
Information about the organizational structure of the Department of Culture of Ivano-Frankivsk City Council, defined by an organizational and administrative document titled 'Structure and staffing'. The dataset originates from the States site of Ukraine and was last updated on January 29, 2025. It is available in common spreadsheet formats including Excel and CSV.
24,495 line images of Arabic handwriting paired with ground-truth text transcriptions. The collection features manuscripts spanning the 19th to 21st centuries, providing a longitudinal view of Arabic script evolution.
10,987 feet of drilling data from the Utah FORGE geothermal well 16A(78)-32, completed 60 days ahead of schedule. The dataset includes survey data, daily operations reports, and rig photos from the 60-day drilling period between October 2020 and January 2021. It documents a highly deviated well reaching a true vertical depth of 8,559 feet in crystalline granite.
District of Columbia's Department of Public Works provides geographic data on waste management infrastructure. The dataset includes locations of street and alley collection points, specific routes, and scheduled collection days. It was last updated in February 2025.
Information about the organizational structure of the Executive Committee of the Black Sea City Council. The dataset was published by the States site of Ukraine and last updated on February 5, 2025. It is available in CSV format.
Image dataset used to fine-tune HoJ, a Stable Diffusion model. The dataset contains between 1,000 and 10,000 samples, is gated to prevent mass downloads, and was authored by n-Arno.
1099 posts from the r/roastme subreddit were scraped for comments before February 27, 2025. The dataset organizes top-level roast comments alongside image analysis for attributes like age, sex, race, and foreground objects. It was created by gus-gustavo and last updated on Hugging Face on 2025-02-27.
Magpie-Align released this dataset in January 2025, providing 250,000 synthetic reasoning records for large language model alignment. It utilizes the Magpie method to generate Chain-of-Thought (CoT) data from DeepSeek-R1 and Llama-3-70B without requiring manual prompt engineering.
A collection of 250,000 instruction-following examples for aligning large language models, created by Magpie-Align. The dataset is derived from the Llama-3-Instruct model's outputs and focuses on reasoning tasks. It was released in 2024, with the dataset page last updated on 2025-01-27.
Arocrbench Doclaynet is a dataset for document layout analysis, created by ahmedheakl and last updated on February 24, 2025. It is associated with the KITAB-Bench project and a corresponding research paper. The dataset likely contains annotated document images for tasks like layout segmentation and object detection.
OCRBench v2 is a benchmark dataset for evaluating large multimodal models, containing at least 10,000 test samples. It includes subsets for English and Chinese, with the English subset containing 7,400 samples. The dataset focuses on visual text localization and reasoning tasks.
Baikovetska village council's organizational structure is documented in this dataset from the States site of Ukraine. The data is available in JSON format and was last updated on February 5, 2025. The specific number of records and detailed fields are not provided in the metadata.
Numerical simulation outputs and validation measurements analyze drifting and blowing snow deposition around HELIOPLANT® Alpine PV structures. The dataset includes results from a sensitivity analysis of key structural parameters and field data from the Gondosolar test-site. Data was produced by ENVIDAT for a related publication and was last updated in 2025.
Vggsound is an audio-visual dataset hosted on HuggingFace by user lenghanz. The dataset was last updated on March 10, 2025, but its specific contents and scale are not detailed in the available metadata. Columns and sample data are unknown, requiring verification after download to understand its full scope.
A directory of enterprises, institutions, and territorial bodies under a specific information manager in Ukraine, including their identification codes in the Unified State Register. The data likely contains contact details such as email addresses, official websites, telephone numbers, and addresses. It was published by the States site of Ukraine and last updated on January 30, 2025.
State-owned enterprises belonging to the sphere of management of the State Service of Ukraine for Geodesy, Cartography and Cadastre are listed. The dataset likely contains identification codes from the Unified State Register of Legal Entities, Individuals Entrepreneurs and Public Organizations. It was last updated on January 30, 2025.
ChatRex is a Multimodal Large Language Model (MLLM) designed to integrate fine-grained object perception with language understanding. The dataset, created by IDEA-Research, supports this model with a decoupled, retrieval-based architecture for object detection using high-resolution visual inputs. It was last updated on January 23, 2025.
Magpie-Align's 150,000-instruction dataset is designed for aligning large language models, specifically for improving reasoning capabilities. It contains synthetic chain-of-thought data generated using models like DeepSeek-R1 and Llama-70B. The dataset was published in a June 2024 technical report and last updated on the Hugging Face platform in January 2025.