Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
A dataset from Kaggle related to reinforcement learning (RL) for the Qwen2.5 Vision-Language Model (VLM). The dataset's title suggests it involves staged code, likely pertaining to training procedures or generated outputs. The specific content, scale, and authorship require verification after download.
SearchVLM is a dataset published on Kaggle. The title suggests it relates to vision-language models, likely containing data for search and retrieval tasks. Specific details on size, creator, and temporal coverage are not provided in the available metadata.
CURATED_VLM_DATASETS_987486 is a dataset collection published on Kaggle. Its title suggests it contains data for training and evaluating Vision-Language Models. The specific contents, size, and origin are not detailed in the provided metadata.
Digital heritage data focuses on the preservation of cultural performances and traditions. The dataset's size, author, and last update date are not specified. It is hosted on the Kaggle platform.
Kaggle hosts this dataset titled 'blipcaptionsoutput'. The title suggests it contains image captions generated by the BLIP (Bootstrapping Language-Image Pre-training) model. The dataset's scale, origin, and specific content are not detailed in the provided metadata.
Kaggle hosts the MedVQA-GI-2026 dataset. It is a multimodal dataset for medical visual question answering, specifically focused on gastrointestinal topics. The dataset's author, organization, and specific scale are not provided in the metadata.
Puffin-4M is a large-scale, high-quality dataset containing 4 million samples for camera-centric multimodal understanding and generation. It integrates vision, language, and camera modalities to address the scarcity of benchmarks in spatial multimodal intelligence. The dataset was created by KangLiao and was last updated in January 2026.
Nemotron-RL-instruction_following combines prompts from the WildChat-1M dataset with verifiable instructions from the Open-Instruct code base. Created by NVIDIA, this dataset is designed for training and evaluating models on objective instruction adherence. It was last updated in January 2026.
TAOBAO-MM is a large-scale recommendation dataset derived from user interaction logs on Taobao, one of the world's largest e-commerce platforms. It features historical behavior sequences of up to 1,000 interactions per user and includes high-quality multimodal embeddings. The dataset was authored by TaoBao-MM and was last updated on the Hugging Face platform on 2026-01-15.
ActionDetectionDatasetVLM is a dataset published on Kaggle. Its title suggests it contains video data annotated for action detection tasks, likely intended for training or evaluating vision-language models. The dataset's specific content, size, and origin require verification after download.
Kaggle hosts the synthvision_medical_vqa dataset, which likely contains synthetic medical images paired with questions and answers for visual question answering tasks. The dataset's author, organization, and specific scale are unknown. Its last update date is also unspecified.
LLaVA-2 is a dataset hosted on Kaggle, likely related to vision-language tasks and multimodal AI. Its specific content, scale, and creation details are not provided in the available metadata. The dataset appears to be intended for training or benchmarking large language models with visual capabilities.
Kaggle hosts the LLaVA-3 dataset, a resource for multimodal AI development. The dataset likely contains paired image and text data for training vision-language models. Its specific size, creator, and update history are not detailed in the provided metadata.
Spa3R Vlm is a dataset for vision-language model tasks, hosted on HuggingFace by the author hustvl. The dataset was last updated on March 6, 2026.
SPRITE is a spatial reasoning dataset for Vision-Language Models (VLMs) developed by zhihelu and released in early 2026. It provides image-text pairs designed to improve embodied intelligence by balancing linguistic diversity with computational precision, as detailed in Arxiv paper 2512.16237.
4 distinct subsets including MSCOCO and VisualNews provide multimodal queries and documents for cross-modal retrieval evaluation. The dataset utilizes queries.jsonl files to benchmark performance on text-only, image-only, and combined image-text search tasks.
BLIP2-OPT-27B is a large-scale vision-language model likely designed for tasks like image captioning and visual question answering. The dataset appears to be hosted on Kaggle, but its specific contents, such as training data or model weights, are not detailed in the provided metadata. Further inspection is required to confirm the exact data format and scope.
Data Nanovlm is a dataset published on the Hugging Face platform by the author LMMs-Lab-Speedrun. The dataset was last updated on February 27, 2026. Its specific content and scale are not detailed in the available metadata.
WavLM-Large is a model for speech representation learning, published on Kaggle. The dataset's specific content, size, and origin require verification after download.
LIMO_VQA is a dataset for Visual Question Answering (VQA) tasks, likely containing pairs of images and associated questions. The dataset is hosted on Kaggle, a popular platform for data science competitions and projects. Specific details on its size, creation date, and authors are not provided in the available metadata.