Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
A large-scale collection of astronomical images paired with descriptive captions and synthetic question-answer pairs, designed for training visual language models. The dataset was created by UniverseTBD and last updated on July 28, 2025. It combines imagery from NASA's Astronomy Picture of the Day, the European Southern Observatory's public archive, and ESA's Hubble Space Telescope.
Iconclass is a classification system for art and iconography. The dataset likely contains structured codes and descriptions for visual symbols and themes. It was published on HuggingFace by davanstrien and last updated on September 10, 2025.
MSR-VTT contains 10,000 video clips paired with 200,000 descriptive captions. The dataset, originally created by Microsoft Research, is a standard benchmark for text-video retrieval and captioning tasks. It was last updated on the platform in August 2025.
Cambrian Vision-Centric Benchmark (CV-Bench) is a dataset introduced in the Cambrian-1 research paper for evaluating vision-centric multimodal large language models. The dataset contains annotations and images pre-loaded for processing with Hugging Face Datasets. It was created by nyu-visionx and last updated on July 20, 2025.
Document Haystack is a benchmark dataset for evaluating multimodal Large Language Models on long-context image and document understanding tasks. It was created by AmazonScience for a 2025 research paper to address the lack of suitable benchmarks for processing long documents. The specific row count, column count, and data size are not provided in the input.
FragFake is a dataset for edited-image detection using Vision-Language Models (VLMs). It contains four groups of examples—Gemini-IG, GoT, MagicBrush, and UltraEdit—each with two difficulty levels: easy and hard. The dataset was created by Vincent-HKUSTGZ and was last updated on July 31, 2025.
K-LLaVA-W is a Korean adaptation of the LLaVA-Bench-in-the-wild, designed for evaluating vision-language models. The benchmark was created by translating the original English text into Korean and reviewing its naturalness through human inspection. It was published by NCSOFT and last updated on July 25, 2025.
OpenGVLab's Doc-750K dataset, referenced in the paper 'Docopilot: Improving Multimodal Models for Document-Level Understanding', is a collection of documents for training AI models. The dataset was last updated on July 22, 2025. It appears to contain a large number of document images, as suggested by unzipping instructions for image archives.
NayanaBench is a multilingual visual question answering dataset designed for evaluating multimodal AI systems. It includes 200 examples each for 22 languages, combining optical character recognition and layout analysis. The dataset was created by Nayana-cognitivelab and was last updated on July 28, 2025.
Over 9.3 million synthetically generated image-text pairs form this multimodal dataset created for training the SmolDocling model. The dataset covers code snippets from 56 different programming languages, with text sourced from permissively licensed sources and images generated at 120 DPI using LaTeX and Pygments. It was created by the docling-project and last updated on July 16, -2025.
3,192 image–annotation pairs form the CalliBench dataset for evaluating Vision Language Models on Chinese calligraphy. The dataset, created by author gtang666, includes tasks for full-page recognition and contextual visual question answering. It was last updated on Hugging Face in July 2025.
Aggregating multiple benchmarks for table understanding, this repository by esborisova was updated in September 2025. It categorizes resources into tasks such as table structure recognition, table-to-text, and table question answering.
A collection of question-answering datasets, including Geo170K, Visualpuzzles, TQA, AI2D, RL, LMMS, ScienceQA, and OK-VQA, uploaded by author GY2233 to Hugging Face on 2025-09-02. The title suggests it aggregates multiple established benchmarks for visual and textual reasoning. The specific content, scale, and structure of the combined data require verification after download.
10,000 to 100,000 multimodal records for cold-start supervised fine-tuning (SFT) in reasoning tasks, released by WaltonFuture in 2025. It supports the research paper 'Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start' by providing initial training data for a two-stage reinforcement learning pipeline.
Three categories of multimodal geo-spatial data—tabular grids, heatmaps, and geographic visualizations—designed for foundation model evaluation. The benchmark tests the ability to process dense numerical values and interpret spatial-temporal dependencies within these grid structures.
The MDocAgent dataset supports a framework for multi-modal document understanding, as described in the associated arXiv paper. The dataset was created by Lillianwei and last updated on August 22, 2025. It is hosted on Hugging Face and is associated with a GitHub repository containing the framework's code.
A large-scale, multimodal dataset for Continuous Bangla Sign Language (BdSL) recognition and translation. It includes video samples of real-life continuous sign language performances paired with gloss sentences representing the signed content. The dataset was created by 'banglagov' and was last updated on July 22, 2025.
Indian Cartoon Blip is a dataset uploaded by Surbhipatil to the Hugging Face platform. The dataset was last updated on 2025-09-02 10:39:41. Its specific content, size, and structure are not detailed in the available metadata.
Facecaption 1M is a dataset of 1 million facial image-text pairs, as indicated by its title. The dataset was created by authors from OpenFace-CQUPT and published in a 2024 arXiv paper. The dataset listing on HuggingFace was last updated on August 1, 2025.
A dataset for mixed-modal instruction tuning created by researchers at the University of California, Los Angeles. It is designed for training biomedical assistants by integrating multimodal information. The dataset page was last updated on 2025-07-19.