Loading...
Loading...
Image classification, object detection, segmentation, face recognition, OCR, image generation, video understanding
17,661 datasets
Magpie-Align's 150,000-instruction dataset is designed for aligning large language models, specifically for improving reasoning capabilities. It contains synthetic chain-of-thought data generated using models like DeepSeek-R1 and Llama-70B. The dataset was published in a June 2024 technical report and last updated on the Hugging Face platform in January 2025.
Over 1,200,000 human annotations form one of the largest datasets for aligning text-to-image models. Rapidata collected this data in approximately four days using their Python API, which is described as accessible for large-scale annotation. The dataset was last updated on January 10, 2025.
Five resources comprise a directory for the Department of Education and Science of the Ivano-Frankivsk City Council. The data includes subordinate legal entities, structural subdivisions, and officials for both the department and a related professional development center. The dataset was last updated on 2025-01-09 and originates from the States site of Ukraine.
Vggsound 16K is a dataset of audio-visual clips, likely containing 16,000 entries, published on the Hugging Face platform by author txya900619. The platform tags indicate it contains audio and text modalities, suggesting it may pair sound recordings with descriptive labels. The dataset was last updated on March 7, 2025.
A directory of enterprises, institutions, and organizations managed by the Ukrainian Recovery Agency, including economic societies where the agency manages corporate rights. The directory likely contains identification codes from the Unified State Register of Legal Entities, Individual Entrepreneurs and Public Organizations, along with contact and location details. The dataset was last updated on 2025-01-22 and originates from the States site of Ukraine.
Information on the organizational structure of the Ternopil District Military Administration, sourced from the States site of Ukraine. The dataset was last updated on 2025-01-29 21:19:40.178646.
Department of Youth Policy of Vinnytsia City Council provides a directory of its subordinate and managed entities. The dataset contains three resources listing organizations, organizational units, and official posts. It was last updated on the State site of Ukraine in January 2025.
Holding approximately 200,000 image-text pairs for LaTeX optical character recognition, featuring both printed and synthetic handwritten mathematical formulas. Developed by linxy and updated in December 2024, the dataset aggregates data from Zenodo, CROHME, and custom-built sources to support the LaTeX_OCR project.
USC-PSI-Lab's Humanoid-X dataset supports the paper "Learning from Massive Human Videos for Universal Humanoid Pose Control". It includes released text descriptions, humanoid keypoints, and humanoid actions data, with only part of the human poses data fully released. The dataset was last updated on January 14, 2025.
Magpie Qwen2.5 Pro 1M V0.1 is a collection of instruction data for aligning large language models. It was created by Magpie-Align and released on Hugging Face, with a technical report published on arXiv in June 2024. The dataset aims to address the lack of open alignment data for models with open weights.
Three resources list executive bodies, enterprises, institutions, and organizations under a specific information manager's purview in Ukraine. The data includes separate legal entities, their structural subdivisions, and officials or employees. It was last updated on 2025-01-13 via the States site of Ukraine.
State of Iowa provides an electronic filing system for documents filed with or issued by the Iowa Utilities Commission since January 2, 2009. Information is organized by docket and filing type, including permits, motions, objections, inspections, comments, applications, tariffs, orders, petitions, franchises, certificates, and consumer complaints.
A benchmark comparison dataset for evaluating large language models, published on Hugging Face by author datalab-to. The dataset was last updated on February 28, 2025. Its specific contents and scale require verification after download.
Lists of normative legal acts and acts of individual action adopted by Ukrainian information administrators. The dataset includes draft normative legal acts and is sourced from the States site of Ukraine. It was last updated on January 22, 2025.
Motion-X++ is a large-scale multimodal dataset for 3D whole-body human motion. It includes 2D keypoints for mesh recovery and motion generation, along with SMPL-X annotations differentiating translations and orientations in camera and world coordinate systems. The dataset, created by YuhongZhang, was last updated on January 14, -2025.
Ocrbench V2 is a multimodal dataset for evaluating optical character recognition systems, containing images paired with text. The dataset was created by author 'ling99' and uploaded to Hugging Face in February 2025. Platform tags indicate it contains at least 100,000 data points and includes both image and text modalities.
TAO-Amodal augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects. The dataset also includes modal segmentation masks. It was created by Cheng-Yen (Wesley) Hsieh and was last updated on the Hugging Face platform on January 11, 2025.
Directory of the Department of Transport and Communications of Ivano-Frankivsk City Council contains three resources: a directory of organizations, their structural units, and their officials. The dataset was published on the States site of Ukraine and was last updated on 2025-01-13 08:52:50.604058. It is available in CSV format.
Upstage's TFLOP dataset is a processed version of PubTabNet, curated for table structure recognition research. It contains a filtered subset of document images with erroneous samples removed and includes corresponding Optical Character Recognition inference results. The dataset was last updated on January 23, 2025.
5,195 fully-annotated abdominal CT volumes constitute one of the largest datasets of its kind. The dataset includes annotations for eight anatomical structures: spleen, liver, kidneys, stomach, gallbladder, pancreas, aorta, and IVC. It was created by the CCVL research group at Johns Hopkins University and last updated on January 16, 2025.