Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,941 datasets
A multimodal dataset for cultural reasoning on antique Chinese porcelains, created by SII-Monument-Valley. The dataset is part of the CiQi-Agent project, which aligns visual perception, tool-augmented reasoning, and cultural knowledge. It was last updated on April 1, 2026.
Leak-CURBER is a dataset and code package created for the NeurIPS 2026 Evaluations and Datasets track. It likely contains multimodal data for evaluating tasks related to enzymatic reactions. The dataset was uploaded by an anonymous author on May 7, 2026.
A dataset likely containing multimodal data for training and evaluating models that detect violent behavior. It is hosted on Kaggle, but specific details about its size, creation date, and authorship are not provided. The content and structure must be verified after download.
61,000 fully annotated frames collected for aerial-ground cooperative perception. The dataset integrates synchronized multimodal sensing data and state information from vehicles and UAVs, covering 19 interaction scenarios and 5 weather conditions. It was created by LOTEAT and last updated on Hugging Face in April 2026.
A multimodal dataset titled 'multimodal_closedset_unified' is hosted on Kaggle. The dataset's specific content, size, and structure are not described in the available metadata. Its author, organization, and last update date are unknown.
A dataset titled 'Tulu3 Instruction Following SFT 16K Bucket' is hosted on Kaggle. The title suggests it is likely a collection of instruction-response pairs for supervised fine-tuning of language models. The specific content, size, and creation details are not provided in the available metadata.
The dataset titled 'vlmclip1' is hosted on Kaggle. Its name suggests a connection to vision-language models, likely containing data for training or evaluating systems like CLIP. The specific content, size, and structure require verification after download.
HandVQA is a dataset introduced in a CVPR 2026 paper by researchers from UNIST, University of Aberdeen, University College London, and Fogsphere. It is designed for diagnosing and improving fine-grained spatial reasoning about hands in vision-language models. The dataset page is hosted on Hugging Face by author kcsayem and was last updated on March 30, 2026.
A multimodal dataset designed for training Vision-Language Models to identify trading exhaustion and opportunities. It was created by author SpaceGhost using a Hindsight Mining technique to capture decision snapshots. The dataset was last updated on HuggingFace on 2026-04-10.
Vero-600k is a collection of data for training and evaluating general visual reasoning models, created by researchers at Princeton University's zlab. The dataset supports broad multimodal reasoning tasks across charts, STEM problems, spatial reasoning, and knowledge grounding. It was released in early 2026.
Dataset-multimodal likely contains multiple data types such as images, text, or audio for training integrated AI systems. Published on Kaggle, its specific content and scale are not detailed in the available metadata. The author, organization, and last update date are unknown.
JAMMEval is a curated benchmark collection for evaluating Vision-Language Models on Japanese Visual Question Answering tasks. It refines seven existing Japanese VQA evaluation datasets through two rounds of human annotation to improve reliability. The dataset was created by llm-jp and was last updated in April 2026.
SFT-Dataset is a curated, medium-scale mixture designed to push a base model toward stronger step-by-step reasoning and reliable instruction following. The dataset was created by SeaFill2025 and was last updated on Hugging Face in April 2026. Quantities are chosen to be trainable on modest GPU budgets while keeping signal density high.
Xperience-10M is a large-scale egocentric multimodal dataset of human experience created by ropedia-ai. It is designed for research in embodied AI, robotics, and world models. The dataset was last updated on March 20, 2026.
A collection of 30,000 real-world chart images paired with detailed natural-language captions, intended for chart understanding and image-to-text research. The dataset was created by the 2077AIDataFoundation and was last updated on April 3, 2026.
WavLM Phase 2 S1 is a dataset hosted on Kaggle, likely containing audio data for self-supervised speech representation learning. The specific content, size, and structure are not detailed in the available metadata. Its origin and creation date are unknown.
INDOTABVQA is a benchmark dataset for evaluating Vision-Language Models on cross-lingual table understanding in Bahasa Indonesia document images. The dataset was created by NusaBharat and is associated with a paper accepted at ACL 2026 Findings. It was last updated on the Hugging Face platform on April 9, 2026.
The dataset is derived from the Niphad Grape Leaf Disease Dataset (NGLD), which contains high-quality images of table grape leaves categorized by disease. The original dataset was created by researchers from Symbiosis Institute of Technology and published on Mendeley Data under a CC BY 4.0 license. This version, uploaded by qingwuuu, appears to be adapted for use with visual language models.
Doc InfographicVQA is a dataset hosted on Kaggle. The dataset likely contains infographic images paired with questions and answers to support multimodal AI research. Its specific size, creator, and creation date are not provided in the available metadata.
Doc MP-DocVQA is a dataset for Visual Question Answering on documents, hosted on Kaggle. The dataset likely contains images of documents paired with questions and answers to test machine comprehension. Specific details on size, creation date, and authorship are not provided in the available metadata.