Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,944 datasets
Doc MP-DocVQA is a dataset for Visual Question Answering on documents, hosted on Kaggle. The dataset likely contains images of documents paired with questions and answers to test machine comprehension. Specific details on size, creation date, and authorship are not provided in the available metadata.
DocVQA is a dataset for visual question answering on documents. It is hosted on Kaggle, but detailed metadata such as author, size, and license are not provided. The dataset's content and structure require verification after download.
FlipVQA-85K is a high-fidelity reasoning benchmark curated from a corpus of 544 college-level educational PDF documents, including expert-authored textbooks and exercise sets. The collection spans 11 academic disciplines, primarily in STEM domains where problems involve rigorous and verifiable reasoning processes. It was created by OpenDCAI and last updated on the platform in April 2026.
Vibe Landing Page Arena is a large-scale human preference dataset for evaluating AI-generated landing page design quality. It contains 36,000 pairwise judgments from 3,492 annotators comparing pages generated by four AI tools across 100 prompts and multiple design dimensions. The dataset was created by datapointai and last updated on Hugging Face in April 2026.
Caveman World Knowledge 150K is an instruction dataset containing approximately 150,000 entries for tuning language models. It was created by author Blackbean109 and was last updated in April 2026. The dataset blends factual world knowledge responses with reactions to unknown questions.
CoMM is a high-quality dataset designed to improve the coherence, consistency, and alignment of multimodal content. The dataset was created by author weisuxi and was last updated on 2026-04-24. It sources raw data from diverse origins, focusing on instructional content and visual storytelling.
dataset_vqa is a dataset hosted on Kaggle. Its title suggests it contains data for Visual Question Answering tasks, which involve answering questions about images. The dataset's specific content, size, and origin are not detailed in the provided metadata.
Agent trajectories from PostTrainBench, a benchmark measuring CLI agents' ability to post-train pre-trained LLMs. The dataset was created by aisa-group and last updated on March 16, 2026. Each agent is given a base LLM, an evaluation script, and 10 hours on an NVIDIA H100 80GB GPU to autonomously improve model performance.
CuriaBench is a collection of evaluation datasets for the Curia foundation model, as described in the associated research paper. The datasets were created by the organization 'raidium' and the benchmark repository was last updated on March 31, III. The data is intended to assess the performance of multimodal AI models in radiology.
CT_Bench is a benchmark dataset designed for evaluating multimodal artificial intelligence models in the analysis of computed tomography scans. The dataset likely contains paired medical images and associated clinical or textual data for structured evaluation tasks. Its creation and maintenance details are not provided in the available metadata.
CURA-VLM appears to be a dataset for vision-language model training or evaluation. It is hosted on Kaggle, but no further details about its size, creator, or specific content are provided. The dataset's purpose likely relates to multimodal AI tasks involving both visual and textual data.
A multimodal benchmark dataset for evaluating computer vision and language models. The dataset likely contains images related to computer networks, paired with annotations for model assessment. It is hosted on Kaggle, but detailed metadata about its size, origin, and specific content is not provided.
ChartNet is a large-scale, high-quality multimodal dataset designed for robust chart understanding and reasoning. It contains over one million chart samples, combining geometric visual patterns, structured numerical data, and natural language descriptions. The dataset was created by IBM Granite and was last updated in March 2026.
A dataset titled 'Cot Oracle Convqa Chunked Sonnet' authored by 'ceselder' and published on the HuggingFace platform. The dataset was last updated on 2026-05-11. Its title suggests it likely contains conversational question-answering data, possibly structured for language model training.
A dataset titled 'circuit-vqa-384a' is hosted on Kaggle. The title suggests it likely contains images of electronic circuits paired with questions and answers. The dataset's author, organization, size, and specific contents are unknown and require verification after download.
Fashion images paired with textual descriptions and sentiment labels, published on Kaggle. The dataset likely contains visual and textual data for analyzing consumer sentiment towards fashion items. Metadata is minimal; actual content requires verification after download.
A multimodal dataset capturing 19.8 hours of expert demonstrations across 315 sessions. It includes synchronized RGB-D video, tactile sensing, eye-gaze tracking, pose annotations, and action labels from 21 occupational therapists performing 15 daily caregiving tasks. The dataset was contributed by the EmPRISE Lab at Cornell University and is hosted on AWS Open Data.
KITScenes LongTail is a dataset for end-to-end driving research focusing on long-tail events. It provides multi-view video data, vehicle trajectories, high-level instructions, and detailed reasoning traces. The dataset was created by KIT-MRT and was last updated on Hugging Face in April 2026.
SLAKE is a dataset for medical visual question answering, a task combining image understanding and natural language processing. It was published on Kaggle, though the specific author, organization, and collection details are not provided in the available metadata. The dataset's size, format, and exact composition require verification after download.
nutriderm-stage7-vqa is a dataset hosted on Kaggle. The title suggests it is a multimodal dataset for visual question answering, likely involving images and text. The dataset's specific content, scale, and origin are not detailed in the available metadata.