Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,947 datasets
OptimusKG is a modern biomedical multimodal Label Property Graph (LPG). The dataset was authored by Lucas Vittor and is hosted on the Harvard Dataverse platform, with a last recorded update on April 14, -2026.
Zoo-Bus VQA is a synthetic visual question answering dataset built for spatial reasoning and object-centric grounding. It contains generated scenes with benches, stop signs, people, animals, and a clock object representing a bus. The dataset was created by author aprilavrilivan and last updated on March 22, 2026.
Opus-4.6 Reasoning 3000x filtered dataset provides a Turkish translation of reasoning data for LLM training. The dataset is created by Chan-Y to support instruction-following and alignment tasks in Turkish. It was last updated on March 22, 2026.
MR-RATE-vista-seg contains voxel-wise multi-label brain segmentation maps predicted for center modality brain MRI volumes in native space. The dataset is part of the MR-RATE vision-language foundation model release by author Forithmus, with a last recorded update in March 2026. It is hosted on Hugging Face and includes platform tags for healthcare, radiology, and multimodal tasks.
GroundSet is a large-scale Earth Observation dataset built on 20 cm resolution optical aerial orthophotos and legally verified cadastral vector data from the French national mapping agency (IGN). It is designed to advance fine-grained spatial understanding for multimodal models. The dataset was created by RogerFerrod and was last updated in March 2026.
A dataset titled 'vqatrec-count-anything' is hosted on Kaggle. The dataset's title suggests a focus on visual question answering and object counting tasks. Its specific content, size, and authorship are not detailed in the available metadata.
VQATRec-Assets-Small is a dataset hosted on Kaggle. Its title suggests it contains assets for visual question answering or text recognition tasks. The dataset's specific content, size, and structure are not detailed in the available metadata.
vqatrec-test-nogt is a dataset hosted on Kaggle. Its title suggests a focus on visual question answering, likely containing test data for model evaluation. The dataset's specific content, size, and origin are not detailed in the available metadata.
Benchmark-LGBioVLM appears to be a benchmark dataset for evaluating large language models on biomedical vision-language tasks. The dataset is hosted on Kaggle, but its specific contents, size, and creation details are not provided in the available metadata. Further details about the data volume, creators, and creation date require verification after download.
Anthropic Hh Rlhf Preprocessed is a dataset published on huggingface by TheHassanSaud. The title suggests it contains preprocessed data from Anthropic's 'HH' (Helpful and Harmless) project, likely used for Reinforcement Learning from Human Feedback (RLHF). The dataset was last updated on 2026-04-24 18:40:45.
Kaggle hosts a dataset for Visual Question Answering (VQA), a multimodal task combining computer vision and natural language processing. The dataset likely contains images paired with questions and corresponding answer annotations. Published on Kaggle, its specific size, creator, and update date are unknown.
PersonaVLM is a framework for transforming general-purpose multimodal large language models into personalized assistants. The work, authored by ClareNie, was accepted for presentation at CVPR 2026.
Pre-tokenized `.bin` shards for efficient Assamese large language model training. The dataset is hosted on Kaggle, but the author, organization, and specific scale are unknown. The last update date is also unknown.
NVIDIA released this collection of dataset blends in March 2026 to document the specific data mixtures used for Reinforcement Learning (RL) training of the Nemotron-3-Super-120B-A12B model. The data is organized into six distinct training stages including Reinforcement Learning from Verifiable Rewards (RLVR), Software Engineering (SWE), and Reinforcement Learning from Human Feedback (RLHF).
Kaggle hosts this dataset titled 'vlmDatacalibC'. The dataset likely contains data for calibrating vision-language models. Its specific contents, size, and creation details are not provided in the available metadata.
A dataset titled 'vlmdataCalibCFull' published on Kaggle. The name suggests it is likely related to calibration for vision-language models. No further details on size, origin, or specific content are available from the provided metadata.
BD-HazardVLM-100-Pilot is a dataset hosted on Kaggle, likely designed to evaluate vision-language models on hazard recognition tasks. The dataset's specific content, size, and collection details are not provided in the available metadata. Its title suggests it may contain a pilot collection of 100 multimodal samples for benchmarking.
A multimodal dataset from HuggingFace provides synchronized vision and tactile glove sensor data across distinct tasks. The dataset includes RGB video at 30 Hz and 720p resolution, lossless 16-bit depth streams, monochrome camera views, and per-frame aligned tactile data in Parquet format. It was created by touchtronix and last updated on March 16, 2026.
DECO-50 comprises over 5 million frames of teleoperated data for bimanual dexterous manipulation with tactile sensing. The dataset includes 50 hours of data collected on real dual-arm robots across 4 scenarios and 28 subtasks. It was created by BAAI-Humanoid and was last updated on Hugging Face in February 2026.
Training corpus for GO-GPT, an autoregressive transformer model for Gene Ontology term prediction. It contains proteins annotated with GO terms, InterPro domains, STRING protein-protein interactions, and metadata sourced from UniProt.