Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,939 datasets
PerturbReason is the training dataset for the AROMA model, a multimodal architecture for virtual cell modeling presented at ACL 2026. The dataset integrates textual evidence, graph topology, and protein sequences to predict the effects of genetic perturbations. It was authored by blazerye and last updated on Hugging Face in April 2026.
Datapoint AI collected ~91,000 human ranking labels for text-to-video generation models. The dataset contains rankings for 5 videos per prompt across 3 quality dimensions, as judged by 15 annotators per dimension. It was last updated on Hugging Face in April 2026.
GitHub documentation files totaling over 2,900 structured entries across 15 languages. The collection is optimized for training large language models and retrieval-augmented generation systems. The author, organization, and last update date are unknown.
A multimodal dataset for medical visual question answering, published on Kaggle. The dataset likely contains pairs of medical images and associated textual questions and answers. Specific details on size, source, and creation date are not provided in the available metadata.
Spectra is a multimodal question-answering training dataset designed for vision-language models. It combines graduate-level science questions from TQA and ScienceQA with open-world knowledge questions from OKVQA and science questions across physics, chemistry, math, and biology from AI2D. The dataset was created by Tamalmajumder and was last updated on April 18, 2026.
SuperviseLab provides professional video annotation data for training multimodal AI models. This public sample dataset demonstrates annotation methodology and output quality across diverse video content categories. All visual assets have been abstracted to protect source privacy, and identifiable metadata has been removed.
1,962 FullHD videos with YUV420 encoding and durations of 10-15 seconds form the open part of the MSU compression artifacts dataset. The dataset, developed by deepfakesMSU, includes videos at frame rates of 24, 25, 30, 39, 50, and 60 fps for evaluating video quality metrics. The full description is available on the dataset page, and the dataset was last updated on April 14, 2026.
LLaVA-Med-v1.5-Mistral-7B is a dataset likely containing a model or associated data for a large vision-language model specialized in medical applications. The dataset is hosted on Kaggle, but its specific contents, scale, and creation details are not provided in the available metadata. Columns, sample data, and authorship information are unknown.
A WavLM model for audio processing, published on Kaggle. The dataset likely contains model weights or related artifacts for speech representation learning. Specific details on the model's architecture, training data, and performance are not provided in the available metadata.
DanQing100M is a large-scale Chinese vision-language dataset containing 100 million image-text pairs, totaling 12 terabytes. It was created by researchers including Hengyu Shen, Tiancheng Gu, and others from DeepGlint-AI, using web data from 2024 to 2025. The dataset is intended for vision-language pre-training tasks.
A dataset for the ACC2026 Track2 competition, likely focusing on augmented visual question answering. Published on Kaggle, its specific content, size, and creation details require verification after download. The dataset appears to be designed for tasks involving both visual and textual data.
GTPBD-MM is the first multimodal benchmark for terraced scenes, integrating optical imagery, textual descriptions, and Digital Elevation Model (DEM) data. The dataset provides three levels of annotations: parcel, mask, and boundary. It was created by author wxqzzw and last updated on April 15, 2026.
MMFace-DiT Dataset provides multimodal conditioning data for high-fidelity, controllable face synthesis. The dataset, created by BharathK333, includes spatial elements like masks and sketches paired with VLM-enriched semantic captions. It was accepted to CVPR 2026 and last updated in April 2026.
Kaggle hosts a synthetic dataset for anaemia screening. The data is multimodal, likely containing a combination of data types such as images, text, or tabular records. Its synthetic nature suggests it was generated for research and development purposes, though specific details on size, origin, and creation date are unavailable.
A dataset likely designed for the ACC2026 Track2 competition, focusing on Visual Question Answering (VQA). It is associated with the Qwen model and is published on Kaggle. The specific content, size, and collection details are not provided in the available metadata.
20,000 Bangla news headlines are paired with corresponding images for multimodal classification tasks. The dataset is hosted on Kaggle, but details about its author, organization, and creation date are unknown. Column-level documentation and file formats are also unspecified.
FashionMV is a large-scale dataset for product-level Composed Image Retrieval (CIR) created by yuandaxia. It contains 127,000 products, 472,000 multi-view images, and over 220,000 CIR triplets, built through an automated pipeline leveraging large multimodal models. The dataset was last updated on April 14, 2026.
2,030 memes form a fine-grained extension of the Hateful Memes dataset, annotated for nuanced analysis of harmful content. It was created by nils-herrmann and last updated on 2026-04 08. The dataset introduces annotation dimensions for incivility and intolerance beyond binary hatefulness.
Multimodal per-fire and per-day raster data covering U.S. wildfire spread from 2016 to 2025. The dataset is hosted on Kaggle and appears to be a collaboration between Deepfire and Stanford. It provides daily snapshots of fire progression.
VizWiz-VQA-Grounding is a dataset likely designed for visual question answering tasks. It appears to be hosted on Kaggle, but detailed metadata about its size, structure, and creation details are unavailable. The title suggests it contains images paired with questions and answers, potentially with grounding annotations linking answers to specific image regions.