Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
DARE (Diverse Visual Question Answering with Robustness Evaluation) is a multiple-choice VQA benchmark created by cambridgeltl. It evaluates Vision-Language Model performance across five diverse categories and includes four robustness-oriented evaluations based on variations in prompts, answer options, output format, and the number of correct answers. The validation split contains images, questions, answer options, and correct answers.
218 million image-text pairs comprise the BLIP3-KALE dataset, featuring knowledge-augmented dense captions. It was created by Salesforce and last updated on February 3, 2025. The dataset is designed to combine web-scale knowledge with detailed image descriptions.
AllenAI presents the Tulu 3 8B Preference Mixture, a research artifact for training large language models. The mixture is composed from multiple preference datasets, including reused on-policy and off-policy data. The collection was last updated on February 4, 2025, and is licensed under ODC-BY-1.0, with some portions being non-commercial.
Dasool's visual question answering dataset focuses on butterflies and moths. It is designed to benchmark Vision-Language Models for tasks like fine-grained species identification and ecological reasoning. The dataset was last updated on 2025-02-18.
Holmes-VAD provides video sequences and textual reasoning labels for explainable anomaly detection, released by pipixin321 in 2025. It serves as the official data source for the Holmes-VAD framework, which integrates Multi-modal Large Language Models with surveillance footage. The dataset is distributed under the MIT license.
A dataset containing 50 million entries designed to improve Vision-Language Models' ability to ground semantic concepts in visual features. Created by Salesforce, it was last updated in February 2025. The data supports tasks requiring precise localization of objects and understanding of referring expressions.
400,000 human preference responses from 82,000 unique annotators evaluating text-to-image model outputs. The dataset categorizes feedback into preference, coherence, and alignment metrics for large-scale model ranking.
7 million diverse images sourced from datasets like COYO-700M and MS-COCO 2017, each paired with both a short and a detailed caption. This re-captioned dataset was created by DAMO-NLP-SG for training the VideoLLaMA 3 multimodal foundation model and was last updated in February 2025.
3DSRBench is a manually annotated benchmark for evaluating 3D spatial reasoning in large multimodal models. It contains 2,100 visual question-answering pairs on MS-COCO images and 672 on multi-view synthetic images rendered from HSSD. The dataset was created by author 'ccvl' and was last updated on the Hugging Face platform in February 2025.
A large-scale multimodal instruction tuning dataset for colonoscopy research, comprising over 300,000 colonoscopic images and 128,000 medical captions generated by GPT-4V. The dataset includes 62 categories and is designed to instruct models to execute user-driven tasks interactively. It was created by ai4colonoscopy and last updated on February 4, 2025.
This is the training split for the Massive Multimodal Embedding Benchmark (MMEB), used to train VLM2Vec models as described in an ICLR 2025 paper. It comprises data from 20 out of 36 datasets selected for evaluating multimodal embedding models across 4 meta tasks.
A curated collection of high-quality synthetic Python unit tests derived from two code instruction tuning datasets: CodeFeedback-Filtered-Instruction and the training set of TACO. The dataset was created by author KAKA22 and last updated on 2025-01-20. It was used to train CodeRM-8B, a unit test generation model.
M2KR-Challenge is a multimodal retrieval dataset created by Jingbiao and last updated on February 4, 2025. It contains 6.42k query samples with images and optional text, and a collection of 47.3k textual passages with associated web screenshots. The dataset is designed for image-to-document and image+text-to-document matching tasks.
Psychocounsel Preference is a text dataset for preference learning in psycho-counseling contexts, created by the Psychotherapy-LLM author group. It is designed to unlock large language models' counseling skills, as described in the associated research paper. The dataset was last updated in March 2025.
Presenting a synthetic preference dataset for instruction tuning, developed by the LLM-jp collaborative project in Japan. It is specifically aimed at ensuring the safety and appropriateness of large language model outputs in Japanese. The dataset was last updated on February 2, 2025.
GIFT-Eval Pre-training Datasets contain 4.5 million univariate and multivariate time series totaling 230 billion data points, spanning seven domains and 13 frequencies. The collection, created by Salesforce, is designed for pretraining foundation models and is explicitly aligned with the GIFT-Eval benchmark to avoid data leakage between training and testing splits.
One of the largest human preference datasets for text-to-image models, containing over 1,200,000 human preference annotations. It was collected by Rapidata using their Python API over approximately four days. The dataset was last updated on January 10, 2025.
Recap-DataComp-1B is a large-scale image-text dataset where descriptions have been enhanced using an advanced LLaVA-1.5-LLaMA3-8B model. The dataset was created by UCSC-VLAA and was last updated in January 2025.
Llavacot Think is a multimodal dataset containing image-text pairs, categorized as having between 10,000 and 100,000 samples. Created by ahmedheakl, it was last updated in March 2025.
A 700,000-record subset of a larger human-annotated dataset for evaluating AI image generation models, split into three annotation modalities. The dataset was created by Rapidata and last updated on January 10, 2025. It is part of a collection that includes separate datasets for coherence and text-to-image alignment.