Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,944 datasets
Encyclopedic-VQA is a visual question answering dataset converted to a unified Parquet schema. The dataset, originally from Google and presented at ICCV 2023 by Mensink et al., contains questions about detailed properties of fine-grained categories. The data is hosted on Hugging Face by the author reonokiy and was last updated on April 1, 2026.
LLaVA-LoRA-Sidewalk is a dataset hosted on Kaggle. The title suggests it contains multimodal data, likely images and text, related to sidewalk environments. Its specific content, scale, and origin require verification after download.
464,044 co-registered image-text pairs from Sentinel-1 and Sentinel-2 satellites form this large-scale dataset. It was created by BIFOLD-BigEarthNetv2-0 to advance vision-language learning for remote sensing data. The dataset was last updated on the platform in April 2026.
SALMUBench is the official evaluation dataset for a CVPR 2026 benchmark on multimodal unlearning. The dataset, authored by cvc-mmu, is designed to assess methods for removing sensitive associations from models. It was last updated on March 30, 2026.
Hugging Face hosts the AwaRes training dataset, created by NimrodShabtay1986 and last updated on March 26, 2026. This multimodal dataset supports a spatial-on-demand VLM inference framework designed to process low-resolution images and selectively retrieve high-resolution crops. The associated paper and project page detail the framework's performance benchmarks and efficiency gains.
A dataset for fine-tuning the MedGemma-4B vision-language model for Bengali medical question answering. The repository contains training and testing configurations for models like Qwen2.5-VL-7B and MedGemma-4B. It was created by iiCEMAN and last updated on April 8, 2026.
MultiNativQA is a multilingual question-answering resource spanning 7 languages, including high- to extremely low-resource ones. It covers 9 locations/cities and includes dialect variations for languages like Arabic. The dataset was created by QCRI and was last updated on March 31, 2026.
A dataset published on Kaggle with the title 'HHRLHF Dataset'. The dataset likely contains text-based examples of human preferences or feedback, intended for training or fine-tuning language models. Its specific content, size, and origin require verification after download.
MMOU is a benchmark for evaluating multimodal models on joint audio-visual understanding and reasoning in long and complex real-world videos. The dataset was created by NVIDIA and last updated on March 28, 2026. It is designed to test models on video, speech, sound, music, and long-range temporal context.
HalluBench is a benchmark dataset for evaluating hallucination in vision language models on geospatial imagery. It was created by AuwAuwAuw and last updated on 2026-04-05. The dataset covers two application domains: emergency disaster assessment and urban scene understanding.
Rlhf Learn provides resources for enhancing reinforcement learning stability and efficiency. It focuses on advanced algorithms like TRPO, PPO, DPO, GRPO, DAPO, and GSPO for optimized policy training. The repository was authored by Dylsimple60 and last updated on 2026-05-19.
CoVAND provides annotations for a negation-aware visual grounding dataset built upon the Flickr30k corpus. The dataset was created by author 2na-97 to support the ICLR 2026 paper on negation-aware vision-language models. It was last updated in April 2026.
A multimodal dataset related to the 2018 Camp Fire event. The dataset is hosted on Kaggle, but its specific contents, size, and origin are not detailed in the available metadata. Further inspection after download is required to confirm the data types, volume, and collection methodology.
CT-RATE consists of 10,000 to 100,000 3D chest CT scans paired with corresponding radiology reports, released by Ibrahim Hamamci in 2024. This multimodal dataset facilitates the development of 3D medical foundation models through vision-language alignment. It supports diverse tasks including visual question answering, image-to-text generation, and zero-shot classification.
Moellava-package is a dataset hosted on Kaggle. The title suggests it likely contains components or data related to a multimodal large language model. Metadata is minimal; actual content requires verification after download.
A dataset for Visual Question Answering (VQA) tasks, likely containing pairs of images and corresponding questions with answers. The title suggests the data may be organized by specific classes or categories. It is published on the Kaggle platform, but the original author, collection date, and dataset size are unknown.
2,004 high-quality AI voice samples derived from a larger collection of approximately 32,000 samples. The dataset was created by LAION through a process of quality filtering and speaker deduplication using speaker embeddings and clustering. It was last updated on March 17, 2026.
InsightVQA is a large-scale benchmark for hierarchical visual question answering that connects emotion understanding with cognitive reasoning. The dataset, created by ziyul707 and last updated in April 2026, is designed to evaluate model capabilities in interpreting emotional causes, grounding evidence, and performing reasoning.
VLM Voice Commands is a text dataset of 50,000 curated natural language commands for Vision-Language-Model robot control. The dataset, created by cagataydev and last updated on 2026-03-22, contains diverse commands covering 10 categories of embodied human-robot interaction.
OSWorld-Verified Model Trajectories contains between 100,000 and 1,000,000 evaluation records of multimodal AI agents performing tasks in real computer environments. Created by xlangai and updated in March 2026, the data captures verified execution paths and screenshots from state-of-the-art models tested on the OSWorld benchmark.