Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Rbyte provides multimodal datasets for spatial intelligence and robotics, released by yaak-ai and updated in February 2026. The collection utilizes MCAP and TensorDict formats to facilitate high-performance spatial computing and integration with PyTorch and Polars.
UNO-Bench is a unified benchmark for exploring compositional relationships between uni-modal and omni-modal capabilities in AI models. The dataset was created by meituan-longcat and was last updated on December 4, 2025. It is accompanied by released evaluation scripts and a scoring model named UNO-Scorer-Qwen3-14B.
A dataset titled 'Multimodal1' published on Kaggle. The title suggests it contains multiple data modalities, such as text, images, or audio, likely intended for AI model training. The author, organization, size, and specific content are unknown.
Kaggle hosts a dataset titled 'multimodal', which likely contains data from multiple modalities such as text, images, or audio for machine learning tasks. The dataset's specific content, size, and creator are not detailed in the available metadata. Its last update date and other descriptive details are unknown.
A dataset for Visual Question Answering tasks, likely containing pairs of images and questions with corresponding answers. It is hosted on Kaggle. The specific size, creation date, and authorship are unknown.
MCD-rPPG is a large-scale multimodal dataset designed for remote photoplethysmography and health biomarker estimation from video. The dataset includes synchronized video recordings from multiple cameras, as described in the paper "Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation". Author wengziheng uploaded the dataset to the Hugging Face Hub, with a last recorded update on 2025-12-09.
MathVision-Wild provides 1,000 to 10,000 photographic versions of the MathVision test dataset captured in diverse physical environments. Created by MathLLMs and updated in late 2025, it transitions digital math problems into real-world visual contexts to evaluate Vision Language Model (VLM) performance.
MinishLab released Semhash in January 2026 to provide a framework for fast multimodal semantic deduplication and filtering. The project utilizes model2vec and vicinity-based hashing to identify near-duplicate records across text and image datasets.
LLaVA-OneVision-1.5-Instruct is a 22 million instruction dataset curated by MVP-Lab for training large multimodal models. It was developed to support the LLaVA-OneVision-1.5 model family and was last updated in November 2025.
DAD-3DHeads provides dense 3D annotations for head alignment and reconstruction from single images, published by PinataFarms for CVPR 2022. The data includes FLAME model parameters and 3D landmark coordinates for 3D Morphable Model (3DMM) fitting. It was developed to address the lack of diverse head poses in existing 2D landmark datasets.
AllenAI provides a dataset for visual question answering tasks. It contains image-text pairs designed for evaluating multimodal language models. The dataset was updated in January 2026.
Emo-CFG is a dataset for emotion-centric video foundation models, accepted at the NeurIPS 2025 conference. It was created by researchers from Nankai University, Pengcheng Laboratory, and Kuaishou Technology. The dataset was last updated on December 7, 2025.
Gemini 3 Pro benchmark dataset for multimodal evaluation. The dataset was created by AliMertTemizsoy and published on Hugging Face in January 2026. It contains image-text pairs for visual question answering tasks.
265,016 images from MS COCO are paired with 1,105,904 questions and 11,059,040 ground-truth answers. The dataset is structured into balanced pairs where each question is associated with two similar images that result in different answers to minimize language bias.
MultiPriv is a dataset of Personally Identifiable Information entities and prompts designed for privacy risk research in large language models. It was created by author CyberChangAn and last updated on December 1, 2025. The dataset is multilingual and multimodal, though attribute-level VLM images are not directly included in the repository due to open-source certificate limitations.
MCD-rPPG is a large-scale multimodal dataset for remote photoplethysmography and health biomarker estimation from video. The dataset contains synchronized video recordings from multiple camera views, designed for the paper 'Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation'.
SToCorpus-88M is a pre-training dataset used for the SToFM multi-scale foundation model for spatial transcriptomics. The dataset is associated with a research paper and model code published on GitHub. Specific details on data volume, structure, and features are not provided in the input.
Salesforce developed UniDoc-Bench in 2024 as a benchmark for multimodal retrieval-augmented generation (MM-RAG). It contains 1,700+ multimodal QA pairs derived from a corpus of 70,000 real-world PDF pages across eight domains. The data links evidence across text, tables, and figures to support complex document-based reasoning tasks.
MindCube is a benchmark for evaluating Vision Language Models' ability to form spatial mental models from limited visual information. It contains 21,154 questions across 3,268 images, created by MLL-Lab. The dataset was last updated in November 2025.
Released by mvp-lab in 2025, this 85-million record multimodal collection supports the mid-training phase of the LLaVA-OneVision-1.5 framework. It aggregates image-text data from eight major sources including ImageNet-21k, LAIONCN, and SA-1B to facilitate democratized multimodal model training.