Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
26,260 science questions paired with 6,206 images sourced from CK-12 Foundation's open educational resources. The dataset includes both text-only and diagram-based visual reasoning questions for middle school science. It was uploaded by 'notefill' to HuggingFace and last updated on 2025-11-21.
Critic-10K provides approximately 10,000 image triplets designed to train models to rectify inconsistencies in AI-generated visual content. Created by ziheng1234 and associated with the 2025 research paper 'The Consistency Critic', the data uses VLM-based selection to pair reference images with degraded and target versions.
REFED is an affective brain-computer interface dataset integrating multimodal brain signals and real-time dynamic emotion annotation. The dataset was created by REFED2025 and last updated on the platform in November 2025. It synchronizes EEG and fNIRS signals to study the neural mechanisms of emotional dynamic evolution.
Facebook introduces AdvancedIF, a benchmark featuring over 1,600 prompts designed to assess large language models. The dataset includes expert-curated rubrics to evaluate proficiency in complex instruction following, multi-turn interactions, and system prompt steerability. It was last updated on November 26, 2025.
Multimodal recordings of candidate interview responses categorized by personality traits and professional performance metrics. This dataset facilitates research in affective computing and automated soft-skill evaluation within human resources contexts by providing synchronized behavioral data.
DEJIMA is a large-scale Japanese multimodal dataset containing 3.88 million image-caption pairs and 3.88 million image-question-answer pairs. It was created by MIL-UT using a reproducible pipeline involving web-scale image collection, strict filtering, evidence extraction, and LLM-based annotation under grounding constraints. The dataset was last updated on December 2, 2025.
4,000 multimodal instruction-tuning samples designed to instill Evidence-of-Thought (EoT) reasoning into Vision-Language Models for remote sensing. The dataset utilizes a Socratic questioning approach to guide models through logical, step-by-step interpretation of satellite and aerial imagery.
VectorInstitute released VLDBench in January 2026 as a large-scale benchmark for evaluating Vision-Language Models (VLMs) and Large Language Models (LLMs) on multimodal disinformation detection. The framework provides a testing ground for AI safety by presenting models with deceptive content that integrates both visual and textual modalities.
SafeVid-350K is a large-scale dataset containing 350,000 preference pairs designed to instill Helpful, Honest, Harmless principles in Video Large Multimodal Models. It covers 30 scene categories and 29 fine-grained safety sub-dimensions. The dataset was created by yxwang and was last updated on Hugging Face in November 2025.
A dataset named 'Llava Onevision 1.5 Rl Data' published on the Hugging Face platform by author mvp-lab. The dataset was last updated on 2026-01-06. Platform tags indicate it contains both image and text modalities, suggesting it is likely a multimodal dataset for training or fine-tuning vision-language models.
Published on HuggingFace by author mm-eval, with a last update timestamp of 2026-01-12 07:15:59. The dataset's title suggests it is a toolkit for evaluating vision-language models. Its specific content, scale, and data types require verification after download.
66,000 human-annotated audio samples of spoken mathematical equations and sentences in English and Russian form the Speech2LaTeX dataset. It is the first fully open-source large-scale dataset for converting spoken math to LaTeX, drawn from diverse scientific domains. The dataset was created by marsianin500 and last updated on November 16, 2025.
Formosa Vision is an open-source visual language dataset focused on Taiwanese local culture, containing over two thousand images selected from the National Cultural Memory Bank 2.0. The dataset was created by the Twinkle AI community using a hybrid method where visual language models generated image dialogues, which were then manually checked and revised by participants. It was last updated on November 20, 2025.
1,885 curated geometric problems across plane, spatial, and solid geometry categories form this benchmark. Each problem includes structured textual descriptions and visual diagrams for multimodal understanding. The dataset, created by OpenRaiser and updated in November 2025, leverages the Lean 4 proof assistant for formal representation.
MathCanvas-Edit contains 5.2 million step-by-step editing trajectories for mathematical images. The dataset was created by author shiwk24 and was last updated on the Hugging Face platform in November 2025. It forms a core component of the MathCanvas framework for training large multimodal models.
WorldCuisines is a massive-scale benchmark for multilingual and multicultural visual question answering focused on global cuisines. The associated paper was accepted to NAACL 2025 and received the Best Theme Paper award. The dataset was last updated on November 14, 2025.
A dataset for training vision-language models, created by NVIDIA. The dataset page includes a version history with updates from August to September 2025. The dataset was last updated on the platform on 2025-10-22.
GroundCUA is a large dataset of real UI screenshots paired with structured annotations for building multimodal computer use agents. It covers 87 software platforms across productivity tools, browsers, creative tools, communication apps, development environments, and system utilities. The dataset was created by Fhrozen and last updated on Hugging Face in November 2025.
10,000 spatial reasoning samples designed for geometric imagination from limited 2D visual perspectives. The dataset facilitates 3D mental modeling during reasoning tasks without the need for explicit 3D prior inputs or depth data.
UniBiomed is a foundation model designed for grounded biomedical image interpretation. The model was created by Luffy503 and was last updated on November 11, 2025. It is based on the MedTrinity dataset, which must be downloaded separately.