Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Afri-MCQA is a multimodal cultural question-answering benchmark. It contains 8,000 Q&A pairs across 16 African languages from 13 countries, created by native speakers. The dataset was published by Atnafu and last updated in January 2026.
Kaggle hosts the VQA Zewail dataset, likely focused on visual question answering tasks. The dataset's specific content, size, and origin are not detailed in the provided metadata. Its creation date and last update are unknown.
A subset of the dataset introduced in the paper 'ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding'. This dataset is designed to train multimodal models for streaming video understanding, focusing on proactive interaction tasks. It was authored by EurekaTian and last updated on the Hugging Face platform in January 2026.
OmniSpatial is a benchmark dataset for evaluating spatial reasoning in vision-language models, as presented in an ICLR 2026 paper. The data is structured in a JSON schema with components like 'id' for question identification. The dataset was created by author 'qizekun' and last updated on January 27, -2026.
A dataset for instruction tuning, likely containing text prompts and responses in the Maithili language. It was published on the Hugging Face platform by the author Bansal123 and was last updated on March 1, 2026. The specific content, size, and collection methodology are not detailed in the available metadata.
MMAU provides between 1,000 and 10,000 test records for evaluating audio large language models, released by TwinkStart in early 2026. It is integrated into the UltraEval-Audio framework to benchmark performance across 12 task types and 10 languages. The data spans four specialized domains: speech, general sound, medical audio, and music.
WAVLM Base Local is a self-supervised speech representation model. It is hosted on the Kaggle platform, but the dataset's specific contents, size, and creation details are not provided in the available metadata. The model's architecture and training methodology are likely detailed in its associated research publication.
Tagavlm Dataset is a multimodal dataset hosted by HuggingFace, created by user tiredtony. It is intended for vision-language model training and was last updated in March 2026. Its specific contents and size are not detailed.
1,000 historical recipes prepared for Vision-Language Model training. The dataset includes JSON metadata, suggesting structured information about the recipes. It is hosted on Kaggle, but the original source and collection methodology are not detailed in the provided metadata.
City of Austin data details the development of a new pedestrian and bicycle bridge over Lady Bird Lake near Longhorn Dam. The dataset is tagged for urban planning, geospatial analysis, and multimodal infrastructure within Austin, United States. It was last updated in March 2026.
3MDAD is a multimodal, multiview, and multispectral dataset focused on driver actions and distraction. It contains video and image data from multiple camera perspectives and spectral bands for analyzing driver behavior. The dataset was created for research in automotive safety and computer vision.
Urban Friction Atlas is a multimodal dataset designed for place suitability prediction tasks. The dataset integrates multiple data types, as indicated by its platform tags, to model urban environments. The author, organization, and specific temporal coverage are not provided.
A subset of the BLIP3o-Pretrain-Long-Caption and BLIP3o-Pretrain-Short-Caption datasets translated into Turkish. The dataset is intended for training or fine-tuning image-to-text models. It was created by the author 'ituperceptron' and was last updated on January 15, 2026.
Deepchestvqa is a dataset hosted on HuggingFace by author ZiyueWang. The dataset's columns and sample data are unavailable, making its exact content and scale uncertain. It was last updated on March 7, 2026.
A dataset likely designed for Visual Question Answering (VQA) tasks, focusing on salience and conflict within images. It is hosted on Kaggle, but specific details about its size, creation date, and authorship are unknown. The dataset's content and scope require verification after download.
MHAL Dataset Annotations for LLaVA is a dataset published on Kaggle. The title suggests it contains annotations for the LLaVA (Large Language and Vision Assistant) model, likely involving multimodal data linking images and text. The dataset's specific content, size, and authorship are unknown.
A dataset for figure question-answering, synthesized for pre-training models. It was created by researchers including Risa Shionoda and Kuniaki Saito for the AAAI-25 Workshop on Document Understanding and Intelligence. The dataset page was last updated on 2026-01-18.
Pano VQA is a dataset hosted on Hugging Face by the user 'wakinghours', last updated on March 2, 2026. Its title suggests it is designed for Visual Question Answering tasks involving panoramic or wide-field-of-view imagery. The dataset's specific content, scale, and structure require verification after download as metadata is minimal.
Kaggle dataset titled 'data_vlm_diff_ready_40_cmd'. The name suggests a collection of data prepared for vision-language models and diffusion processes. The dataset's specific content, size, and origin are not detailed in the provided metadata.
A multimodal dataset from the LLaVA-CoT project, likely containing image-question-answer pairs structured for visual reasoning tasks. The dataset includes a train.jsonl file with conversation data linking images to questions and answers, suggesting a format for training or evaluating vision-language models. It was authored by 'berhaan' and last updated on 2026-01-17.