Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
RoboVQA contains video and text data for training models to answer questions about robotic scenes. The dataset includes over 100,000 entries, as indicated by its Hugging Face size category. It was created by Tianli and last updated in July 2025.
Giving access to reasoning traces generated by Gemini-2.5-pro for the Robo2VLM-1 visual question answering benchmark. It contains logical, step-by-step explanations that justify correct answers for robotic manipulation tasks across diverse, in-the-wild environments.
BLIP3o-60k is a dataset distilled from GPT-4o for instruction tuning of text-to-image models. It includes categories such as JourneyDB, human-centric data from MSCOCO, Dalle3 outputs, Geneval, common objects, and simple text. The dataset was created by BLIP3o and last updated on May 25, 2025.
Therapeutics Data Commons (TDC) is a collection of multimodal benchmarks and datasets for drug discovery and therapeutic science developed by the Harvard MIMS group. Updated as recently as July 2025, it provides a standardized framework for evaluating machine learning models across the drug development pipeline.
A 2024-09-01 upload of filtered VisualWebInstruct data for the OneVision training stage. The dataset, created by lmms-lab, contains subsets like ureader_kg and ureader_qa, provided as processed JSON files and compressed image folders.
4,992 social media posts from the RedNote platform categorized into 613 advertisement and 4,379 non-advertisement samples. The dataset includes 26,324 associated images distributed across training, validation, and test splits for covert marketing detection.
Vqa Multitask is a dataset for multitask learning, likely combining visual and textual data for question answering. It was published on huggingface by author WaltonFuture and was last updated on July 9, 2025. The specific content, scale, and structure require verification after download.
2 document images from the DocVQA dataset serve as fixtures for the HuggingFace Transformers library. These samples facilitate the testing of LayoutLMv2FeatureExtractor and LayoutLMv2Processor across specific unit test files.
A benchmark dataset comprising over 14,500 questions on non-synthetic images, created to assess stereotype biases in Large Multimodal Models (LMMs). The dataset, authored by ucf-crcv, was last updated on May 16, 2025. It spans nine diverse domains and 54 sub-domains to rigorously evaluate LMM performance in visually grounded stereotypical scenarios.
MSR-VTT is a benchmark dataset for text-video retrieval, containing 10,000 video clips and 200,000 captions. It was introduced in the 2016 paper 'MSR-VTT: A large video description dataset for bridging video and language' and is hosted on Hugging Face by user friedrichor. The dataset uses a standard 1K-A split protocol with training sets of 7,010 and 9,000 videos and a test set of 1,000 videos.
VS-Bench is a multimodal benchmark for evaluating Vision-Language Models in multi-agent environments. The benchmark evaluates fourteen state-of-the-art models across eight vision-grounded environments using two complementary dimensions. It was created by author zelaix and last updated on June 4, 2025.
A training and evaluation corpus for VDocRAG, a retrieval-augmented generation framework designed to understand real-world documents from visual features. The dataset is a unified collection of open-domain document visual question answering data, encompassing diverse document types and formats. It was created by NTT-hil-insight and last updated on 2025-05-26.
2,101 image-text pairs designed for unsupervised post-training of multi-modal large language models. Each entry includes a 'problem' field with a geometric reasoning question and an 'answer' field containing the corresponding solution.
Core-Five is a multi-modal geospatial dataset built for foundation models, unifying Earth Observation data from five essential sensors into aligned spatiotemporal datacubes. It includes optical Sentinel-2 data at 10m resolution and other sensor data for multi-modal vision tasks.
MedTrinity-25M consists of 25 million multimodal medical records featuring multigranular annotations, developed by UCSC-VLAA for ICLR 2025. The dataset provides large-scale image-text pairings designed to advance the training and evaluation of medical Multimodal Large Language Models (MLLMs).
Llava 3D Data is a multimodal dataset published on HuggingFace by author ChaimZhu. The dataset was last updated on July 11, 2025. Its specific content and scale are not detailed in the available metadata.
5 million images are each paired with a short caption generated by the Qwen/Qwen2.5-VL-7B-Instruct model. The dataset was created by BLIP3o and last updated on Hugging Face in May 2025. It is intended for pretraining vision-language models.
4 million images from the JourneyDB collection, hosted by the BLIP3o organization. The dataset was last updated on May 26, 2025. It is intended for use in pretraining multimodal AI models.
A dataset created by nhagar on May 15, 2025, providing the URLs and top-level domains associated with training records in the HuggingFaceFW/fineweb dataset. It was created by downloading source data, extracting URLs and domains, and retaining only those identifiers to make exploring LLM training datasets more accessible.
WildDoc is a dataset created by ByteDance to evaluate the document understanding capabilities of vision-language models in real-world scenarios. It is designed to facilitate the understanding of documents in the wild, as described on its project homepage. The dataset was last updated on May 19, 2025.