Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,936 datasets
WildTableBench is a benchmark dataset for evaluating multimodal foundation models on table understanding in the wild. It contains 402 real-world table images collected from diverse domains and 928 questions across 5 categories and 17 subtypes. The dataset was created by author jzhuang and was last updated on Hugging Face in May 2026.
OPI-Struc is a multimodal instruction-tuning dataset designed for the STELLA project. The dataset was created by BAAI and its related paper was accepted at ACL 2026. The dataset page was last updated on May 12, 2026.
RoboFAC is a multimodal visual question-answering dataset for robotic failure analysis and correction. It comprises over 10,000 robot manipulation videos and 78,623 question-answer pairs, supporting tasks across simulated and real-world environments. The dataset was created by MINT-SJTU.
217 examples across 7 top-level categories and 23 subcategories comprise this benchmark for evaluating multimodal models. Created by zai-org, the dataset requires models to identify entities and perform multi-step reasoning with search-augmented information to answer complex questions. It was last updated on 2026-05-16.
Christopher Mai published per-fold test results for a fine-tuned LLaVA-1.5-7B model on the MVTec zipper dataset. The 5.5 KB dataset contains metrics reported as percentages, except for the Kappa value. It was last updated on April 29, 2026.
BioMatrix-SFT is the supervised fine-tuning corpus used to train the BioMatrix multimodal foundation model. The model integrates 1D sequences, 3D structures, and natural language for molecules and proteins within a single decoder-only architecture. The dataset was created by QizhiPei and was last updated on the Hugging Face platform in May 2026.
Tencent's benchmark evaluates LLM performance on complex translation instructions. It covers 6 constraint types across multiple languages, including single-constraint and multi-constraint scenarios. The dataset was last updated on 2026-05-20.
DiscoverLLM-multiturn-preferences is a dataset of multi-turn dialogue data with scored candidate completions. It was produced by best-of-N synthesis over the DiscoverLLM user simulator and is authored by kixlab. The dataset was last updated on 2026-05-13.
CiteVQA is a document visual question answering benchmark designed to evaluate faithful evidence attribution. The dataset contains 1,897 question-answer pairs grounded in real-world PDF documents. It was created by opendatalab and last updated on 2026-05-13.
OpenStreetCLIP Dataset contains satellite imagery aligned with OpenStreetMap vector metadata for training vision-language models. The dataset is organized into sharded TAR archives for efficient streaming. It was uploaded by alessiopierdominici to Hugging Face and last updated on 2026-05-10.
950 test rows comprise the SalArt-VQA benchmark for visual question answering focused on salient artifacts in AI-generated images. The dataset includes 475 artifact images, 356 clean real-image references, and 119 paired generated artifact-free counterparts. It was created by salartvqa and last updated on Hugging Face in May 2026.
A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. The dataset is associated with a published paper, a hosted challenge, and code, and was last updated on 2026-05-07. It was created by the author 'orena-dkfz'.
NCCE31_Natthapol_Scaffolding_Dataset is a multimodal dataset for research on using foundation models to create construction scaffolding masks for image segmentation. The dataset is 9.4 MB in size and includes JPG and JSON files. It was authored by Natthapol Saovana and last updated on April 24, 2026.
An AI model dataset combining training metrics with arena performance. It was sourced from Kaggle, but the author, organization, and last update date are unknown. The dataset's specific size, row count, and file formats are also unspecified.
KITScenes Multimodal is a high-fidelity autonomous driving dataset designed for research toward production-grade urban driving. It focuses on complex European city environments and combines high-resolution sensor data. The dataset is an early pre-release from KIT-MRT, last updated on May 6, 2026.
WebEyes is a task-level benchmark for evaluating search-based visual reasoning, released by yangbokang81 and last updated on May 13, 2026. It supports three distinct datasets: WebEyes-Ground, WebEyes-Seg, and WebEyes-VQA. Each task is released as a JSONL file, with mirrored Parquet files used for direct image rendering on the Hugging Face platform.
Free-text descriptions of proteinβprotein interactions (PPIs) pairing UniProt accessions with explanatory paragraphs. The dataset was built by xiao-fei to train and evaluate multimodal models that generate PPI descriptions from protein sequence and structure inputs. It was last updated on 2026-05-12.
furproxy provides a collection of captions for furry-themed images sourced from platforms like e621, CivitAI, and booru sites. The dataset contains approximately 7,500 captions, with at least 70% of the complex scenes being human-reviewed and edited. Captions were generated using Gemini 3 Flash and processed through a pipeline involving multi-crop passes and combination.
A subset of Google DeepMind's RoboVQA dataset, re-hosted for loader compatibility. Human-annotated long-horizon robotics video question-answering data across three embodiments, used to train the allenai/Molmo2-ER-4B model. The upstream dataset is described in the paper 'RoboVQA: Multimodal Long-Horizon Reasoning for Robotics' (arXiv:2311.00899).
CMDPAD challenges the static personality assumption by providing dynamic utterance-level scores for the Big Five personality traits. The dataset moves beyond emotion recognition to predict the emotional trajectory of the next interaction turn. It was authored by HensonXie and last updated on Hugging Face in May 2026.