Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Synthetic faceβiris dataset designed for multimodal biometric research and testing. The dataset's author, size, and specific creation details are not provided. Its last update date and licensing terms are also unknown.
CCTV-Pedestrian-1K is a dataset of high-angle surveillance pedestrian images intended for training Vision Transformers (ViT) and Vision-Language Models (VLM). The dataset is hosted on Kaggle and is tagged for applications in public safety and computer vision. Specific details on the number of images, collection time, and creator are not provided in the available metadata.
Multimodal cardiac data integrates electrocardiogram (ECG), photoplethysmogram (PPG), and cardiac timing features. The dataset is hosted on Kaggle and is associated with platform tags for biology, signal processing, and medicine. Specific details on size, origin, and update frequency are not provided in the available metadata.
Aligned text, image, and audio data for cross-language AI translation tasks in Traditional Chinese Medicine (TCM). The dataset is hosted on Kaggle and is tagged as suitable for beginners. Its author, organization, and specific size are unknown.
MGI-TED provides multimodal features for analyzing toddler development and learning behavior. The dataset's author, organization, and specific scale are currently unknown. It is hosted on Kaggle, but details on its collection method and temporal coverage are not provided.
Longtimescope is a dataset for exploring long-video understanding with large multimodal models, as referenced in the Apollo2 research paper. The dataset was created by the Apollo-LMMs team and was last updated on the Hugging Face platform in January 2026. Its specific size, format, and content details are not provided in the available metadata.
NVIDIA released this collection of approximately 9 million vision-language samples in late 2025. It focuses on document understanding, visual question answering, and video-to-text tasks across multiple languages.
S-Chain is a multimodal medical dataset developed by Khai Le-Duc and a multi-institutional research team, last updated in December 2025. It provides structured visual chain-of-thought reasoning paths for clinical tasks across eight languages, including English, Arabic, and Japanese. The data supports a wide range of tasks from object detection to multilingual text generation.
WildfireVLM is a dataset hosted on Kaggle, likely focused on visual and language modeling for wildfire events. The platform tags suggest it contains geospatial and computer vision data, potentially for benchmarking deep learning models. Its specific content, size, and creation details require verification after download.
HAIM Multimodal Full Dataset is hosted on Kaggle. The dataset's specific content, size, and creation details are not provided in the available metadata. Its title suggests it contains multiple data modalities, likely for machine learning research.
Image and text question-answer pairs representing 90 distinct animal species. It provides structured data for Visual Question Answering (VQA) tasks, focusing on the identification and description of fauna.
A synthetic electronic health record dataset integrating text notes and time-series vital sign data. The dataset is designed for healthcare predictive research, specifically HPR. It was created by an unknown author and published on Kaggle, with no information on its size or last update.
HADES-VLM-Data is a dataset for training vision-language models, published on Kaggle. The dataset's specific content, size, and creation details are not described in the available metadata. Its intended use likely involves aligning visual and textual information for AI model development.
A multimodal dataset focused on student engagement, published on Kaggle. The dataset likely contains multiple data types such as video, audio, or sensor readings to capture behavioral and interaction patterns. Specific details on volume, collection method, and authorship are not provided in the available metadata.
Patch embeddings for the CAMELYON16 dataset generated using the UNI foundation model. The embeddings are derived from 128x128 micrometer tissue patches, with segmentation and patching performed using a modified version of the CLAM toolkit. The dataset was authored by kaczmarj and last updated on December 10, 2025.
A collection of question-answer pairs in the Myanmar language designed for instruction tuning of Large Language Models. The dataset aggregates content from multiple sources covering domains like agriculture, health, microbiology, general knowledge, and Buddhism. It was created by chuuhtetnaing and last updated on Hugging Face in December 2025.
JMMMU-Pro is an image-based Japanese multi-discipline multimodal understanding benchmark. It extends the JMMMU benchmark by composing question images and text into a single image, requiring integrated visual-textual understanding. The dataset was created by JMMMU and last updated on Hugging Face in December 2025.
SpecVQA is a benchmark dataset for evaluating Multimodal Large Language Models on spectral understanding and visual question answering tasks using scientific images. The dataset is authored by UniParser and was last updated in December 2025. It contains images and text data, with specific row and column counts unknown.
The AQI Multimodal Dataset is a collection of data related to air quality, likely containing measurements from various sources. The dataset is hosted on Kaggle, but specific details about its size, origin, and creation date are not provided in the available metadata. Further verification is required to confirm the exact contents, scale, and authorship.
598,000 high-quality samples for training and evaluating multimodal code generation models. The dataset covers HTML generation, chart-to-code, image-augmented QA, and algorithmic problems, supporting research in unifying vision-language understanding with code generation. It was created by author 'lingjie23' and last updated on December 24, β2025.