Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
This robotics dataset contains 3,000 episodes and 149,985 frames of multimodal data collected from a Kuka robot arm. Released by the LeRobot team and associated with research paper 1810.10191, the collection provides 20 FPS video and time-series sensor data for a single robotic task.
A dataset introduced in a 2025 paper titled 'Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari'. It supports research in transliterating the ancient Modi script of Maharashtra into the modern Devanagari script used for Marathi and other languages. The dataset was created by author historyHulk and last updated on Hugging Face in September 2025.
APTO-001 developed this dataset to improve large language model instruction-following capabilities. The dataset likely contains synthetic text examples designed to train models on handling complex, multi-step instructions, as described in the platform description. It was last updated on September 12, 2025.
Longitudinal data from 9 to 12 months of age analyzes the role of touch in infant language development. The dataset includes databases for analysis and video clips with examples of categorized behaviors. It was authored by Murillo Sanz, Eva and last updated on October 14, 2025.
A large-scale dataset constructed for medical visual question answering (Med-VQA) tasks. It is based on the ReXGroundingCT data and contains CT volumes paired with multi-class segmentation masks, where each mask channel represents a specific lesion type. The dataset was uploaded by liyf001 and last updated on October 11, 2025.
Misraj Structured Data Dump (MSDD) is a large-scale Arabic multimodal dataset created by Misraj. It was extracted and filtered from Common Crawl dumps using a WASM pipeline and uniquely preserves the structural integrity of web content by providing markdown output. The dataset was last updated on September 29, -2025.
KIE-HVQA is a dataset supporting research on mitigating Optical Character Recognition hallucinations in multimodal large language models. The dataset was created by bytedance-research and is associated with a paper accepted by the NeurIPS 2025 Main Conference. The data likely contains multimodal document samples for evaluating and improving OCR integration in vision-language models.
Caption3o-LongCap-v4 is a large-scale, high-quality image-caption dataset designed for training and evaluating image-to-text models. It is derived from prithivMLmods/blip3o-caption-mini-arrow and additional curated sources, emphasizing long-form captions covering a wide range of real-world and artistic scenes. The dataset was last updated on 2025-09-15 by prithivMLmods.
Caption3o-XL-v4 is a large-scale, high-quality dataset derived from prithivMLmods/blip3o-caption-mini-arrow and other curated sources. It is designed for training and evaluating image-to-text models, with an emphasis on long-form captions covering a wide range of real-world and artistic scenes. The dataset is in Parquet format, contains English text, and was last updated on September 15, 2025.
Approximately 100,000 image-caption pairs form this dataset for training image-to-text models. It was created by prithivMLmods and last updated on August 28, 2025. The dataset emphasizes long-form captions covering a wide range of real-world and artistic scenes.
MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. This mixture contains a subset of OXE formulated as Action Reasoning Data along with auxiliary robot data and a link to Multimodal Web data. The dataset page was last updated on September 10, 2025.
A PyTorch-based implementation of the OpenAI CLIP architecture for image-text alignment, authored by Moein Shariatnia and updated in October 2025. It provides a dual-encoder framework for processing image-text pairs using BERT for natural language processing and Vision Transformer components.
ChemVQA Text is a dataset published on HuggingFace by author chandrabhuma. The title suggests it likely contains chemistry-related content for visual question answering tasks. The dataset was last updated on October 28, 2025.
ShizhenGPT's pre-training dataset contains over 5 billion tokens of Traditional Chinese Medicine text from websites and books, along with a large-scale image-text dataset. The dataset was created by FreedomIntelligence and was last updated in September 2025.
Over 5 billion tokens of Traditional Chinese Medicine text form the largest existing TCM corpus, sourced from websites and books. FreedomIntelligence released this multimodal dataset for pre-training the ShizhenGPT model. It was last updated in September 2025.
Mizzen AI, CUHK MMLab, and academic partners released the Human Preference Dataset v3 (HPDv3) in August 2025. It comprises 1.08 million text-image pairs and 1.17 million annotated pairwise comparisons for modeling human preferences. The dataset is associated with the ICCV 2025 paper 'HPSv3: Towards Wide-Spectrum Human Preference Score'.
MMEB-V2 is a benchmark dataset for evaluating multimodal embedding models, created by VLM2Vec and updated in September 2025. It expands upon a previous version to include five new tasks: Video Retrieval, Moment Retrieval, Video Classification, Video Question Answering, and Visual Document Retrieval.
A multimodal dataset for Point of Interest recommendation based on the Yelp Open Dataset. It includes business metadata, user reviews, business photos, and LLM-generated summaries of reviews and images. The dataset was uploaded by wzehui on September 8, 2025.
550 annotated speech samples categorized across 11 distinct paralinguistic dimensions for speech-to-speech model evaluation. The dataset includes curated audio files and corresponding annotations derived from the Step-Audio 2 technical research.
SimLingo-Data consists of 3,308,315 samples of autonomous driving data generated in the CARLA 2.0 simulator by RenzKa. It integrates sensor readings and action labels with natural language annotations for driving commentary, instruction following, and visual question answering, collected using the PDM-Lite rule-based expert.