Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Rapidata's Flux SD3 MJ Dalle Human Alignment Dataset is one of three splits from a larger collection of over 2 million human annotations for image generation models. This specific subset focuses on text-to-image alignment, while the other splits cover preference and coherence judgments. The dataset was last updated on Hugging Face in January 2025.
PathMMU is a massive multimodal expert-level benchmark for understanding and reasoning in pathology. It was released by author jamessyx on Hugging Face, with the benchmark data and evaluation code published on August 7, 2024. The dataset is intended to address the lack of specialized, high-quality benchmarks for large multimodal models in the pathology domain.
MotionBench is a benchmark dataset designed to evaluate and improve the fine-grained motion comprehension capabilities of vision-language models. The dataset was created by zai-org and released in January 2025. It aims to guide the development of more capable video understanding models.
Aggregating 10,000 to 100,000 medical image-text pairs, this 2024 release from FreedomIntelligence serves as a standardized evaluation suite for multimodal LLMs. It incorporates six distinct benchmarks including VQA-RAD, SLAKE, and PathVQA to test models like HuatuoGPT-Vision.
12.4 million image-caption pairs constitute the largest public domain image-text dataset for training foundation models. The dataset was created by Spawning and released in October 2024, as indicated by the associated arXiv paper identifier. It features community-driven governance mechanisms aimed at reducing harm and supporting reproducibility.
LLM-jp, a collaborative Japanese project, created this synthetic dataset for instruction tuning. It contains a subset of the 801,000-instruction Aratako/Synthetic-JP-EN-Coding-Dataset. The dataset was last updated in January 2025.
PDF-WuKong is a dataset for training and evaluating large multimodal models on long PDF documents. The data accompanies the research paper 'PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling'. Author yh0075 uploaded it to Hugging Face on 2025-01-06.
Multimodal product data from the Rakuten France e-commerce platform, contributed by user yassinemtg and updated in February 2025. The dataset is designed for classification tasks, combining visual and textual information. It focuses on products sold in the French market.
HQD4VLM is a dataset curated for vision-language model research. The dataset likely contains filtered samples intended to reduce noise and improve training efficiency. It was created by author Nhanvi282 and last updated on January 11, 2025.
Vision-language document retrieval training pairs transformed from the vidore/colpali_train_set for Tevatron compatibility. The data is structured to support the training of multi-vector retrieval models like ColPali within the Tevatron ecosystem.
Mantis-Instruct contains 721,000 instruction tuning examples across 14 specialized subsets. It is a fully interleaved text-image dataset designed for training multimodal models on skills like co-reference, reasoning, and temporal understanding. The dataset was created by TIGER-Lab for training the Mantis model families.
8,000 verified multimodal examples for instruction tuning and vision-language tasks, created by the LMMS-Lab. The dataset was last updated in January 2025 and is hosted on Hugging Face.
HH_length_biased_15k is a 15,000-sample subset of Anthropic/hh-rlhf, created for the paper 'Understanding impacts of human feedback via influence functions'. Taywon Min authored this dataset, which was last updated on December 5, 2024. It contains 976 samples where responses were intentionally flipped to be lengthy.
TIGER-Lab's OmniEdit Filtered 1.2M dataset, last updated December 2024, is designed for training a general-purpose image editing model. The dataset was created by filtering data using large multimodal models like GPT-4o for quality assessment. It provides supervision for seven distinct image editing tasks.
23,167,456 Midjourney-generated images and captions were compiled by deepghs and an anonymous provider as of December 2024. This collection contains original image files alongside metadata such as dimensions and unique identifiers.
ChartX & ChartVLM is a benchmark and foundation model designed to evaluate the ability of Multi-modal Large Language Models to query and reason with information from visual charts. The dataset was created by InternScience and was last updated on Hugging Face in December 2024. It is intended to comprehensively and rigorously test chart understanding and reasoning capabilities.
29,980 synthetically generated examples designed to enhance a model's ability to follow instructions precisely and satisfy user constraints. The dataset was curated by the Allen Institute for AI and uses a persona-based methodology to generate diverse instructions, with constraints borrowed from the IFEval taxonomy. It was last updated on November 21, 2024.
This benchmark contains millions of nature photographs paired with expert-level scientific queries for text-to-image retrieval tasks. It evaluates multimodal models on their ability to process complex biological and ecological inquiries against large-scale image collections to support scientific discovery.
633,565 multimodal records of anime, manga, and game characters sourced from 3,860 Fandom wiki sites. The dataset pairs character images with metadata extracted from HTML and descriptive captions generated by the Qwen-VL-72B-Instruct vision-language model.
17,736 galaxies are labeled by citizen scientists through the Galaxy Zoo 2 project. The catalog includes right ascension, declination, redshift, and Galaxy Zoo 2 labels for each entry. Leung et al. published this dataset in 2018, and it is hosted by MultimodalUniverse on Hugging Face.