Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,932 datasets
UMI-VQA-8M is a large-scale visual question answering dataset built for UMI-style wrist-mounted fisheye observations. It contains 8 million visual question-answering samples and provides visual-language supervision for UMI observation scenarios. The dataset was created by TeleEmbodied and was last updated on the Hugging Face platform in June 2026.
GeoMeld is a large-scale multi-modal remote sensing dataset introduced in a CVPRW 2026 paper. It contains approximately 2.5 million spatially aligned samples spanning heterogeneous sensing modalities and resolutions, paired with semantically grounded captions. The dataset was created by author vimageiitb and last updated on 2026-06-04.
prithivMLmods's dataset contains 27,048 English image-caption pairs, with images at 512x512 resolution. The data is derived from curated sources like blip3o-caption-mini-arrow and was last updated on May 17, 2026. It is designed for training and evaluating image-to-text models.
Yezhi Cui's study on figshare, last updated April 22, 2026, investigates the neural mechanisms of Mandarin tone sandhi perception. The dataset includes functional near-infrared spectroscopy (fNIRS) data and behavioral responses from 44 Vietnamese-speaking learners during a tone discrimination task. The data was collected to compare a gesture training group with a no-gesture control group.
MMS-VPR is a large-scale multimodal street-level visual place recognition dataset created by Yiwei-Ou. It comprises 110,529 images and 2,527 video clips with textual annotations, featuring day-night coverage and a 7-year temporal span in dense pedestrian-only environments. The dataset was last updated on HuggingFace on May 20, 2026.
Xuhong Nan published a study on figshare in April 2026 analyzing multimodal ultrasound data from 65 patients with type 2 diabetes mellitus and 27 control subjects. The dataset includes measurements of carotid intima-media thickness, blood flow velocities, wall shear stress, and pulse wave velocity. The study explores subclinical vascular changes associated with diabetes.
AbstractPhil's diffusion-pretrain-set-ft1 is a multi-source image-caption dataset assembled from seven upstream sources via a uniform ingest pipeline. It was built for the finetune-1 stage of the sd15-flow-lune model family but is applicable to any Stable Diffusion 1.x conditioning experiment. The dataset was last updated on May 22, 2026.
~1.51 million samples across eight splits comprise this multimodal reasoning dataset. It was created by RuoliuYang and last updated on May 25, 2026. The splits include text-based chain-of-thought reasoning, bounding box manipulations, and visual representations like depth maps.
DeepTumorVQA v2 is a 3D abdominal-CT diagnostic Visual Question Answering benchmark containing 438,000 total QA pairs. The dataset includes 10,000 curated benchmark pairs and a 428,000-pair training pool, along with pre-extracted 2D and video modalities and 20,000 agent training trajectories. It was created by the tumor-vqa organization and was last updated in May 2026.
ETCHR GRPO-10K is a dataset of 10,000 multimodal samples created by internlm for enhancing model editing capabilities. It contains five specific tasks: Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle, and Spatial Understanding. Each sample includes an image to be edited and an editing instruction.
Raw evaluation metrics and execution telemetry logs from running the Mostly Basic Python Problems (MBPP) benchmark against the Qwen3 8B dense foundation model. The dataset documents zero-shot functional programming synthesis performance under standard local execution bounds. It was authored by ShahzebKhoso and last updated on May 29, 2026.
Wanqing Peng's protocol document outlines a triple-arm randomized controlled trial investigating Qihuang needle therapy for Parkinson's disease. The trial, registered as ITMCTR2025000402, will enroll 69 patients to assess clinical efficacy and neuroplasticity via multimodal MRI. The document was last updated on 2026-04-13.
A clinical trial protocol for a triple-arm randomized controlled trial investigating Qihuang needle therapy for Parkinson's disease. The dataset includes clinical outcomes and multimodal MRI neuroplasticity markers from 69 patients. The protocol was authored by Wanqing Peng and uploaded to figshare in April 2026.
A 2026 protocol document for a randomized controlled trial investigating a novel acupuncture technique for Parkinson's disease. The study, authored by Wanqing Peng, involves 69 patients and uses multimodal MRI to assess neuroplasticity mechanisms alongside clinical symptom scales. The document is a 104.0 KB PDF published under a CC-BY-4.0 license.
A dataset designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following. The chat subset uses human-written prompts from sources like lmarena, lmsys, and wildchat as seed prompts, with responses generated by GLM-5 and selected via pairwise comparisons using a reward model. It was authored by NVIDIA and last updated on the platform in June 2026.
Dog100K is a high-quality dataset containing over 100,000 image-text pairs of dogs. It is designed for image-text retrieval, multimodal learning, and conditional image generation tasks. The dataset was created by choucsan and last updated on Hugging Face in May 2026.
1.1 GB of data supporting a method for automated long-term tracking of Antarctic ice shelf rift propagation. The dataset, authored by Zixiao Guo and last updated in May 2026, is shared under a CC-BY-4.0 license. It likely contains spatiotemporal corrections and tracking results derived from satellite imagery.
ArtiFact is a large-scale multimodal benchmark combining artwork records from the Rijksmuseum, the Metropolitan Museum of Art, and the Art Institute of Chicago. The dataset contains aligned images and structured metadata, normalized across fields for artists, dates, materials, and techniques. It was created by deem-data and last updated on June 8, 2026.
PP2-M is a multimodal dataset derived from Place Pulse 2.0, enriched with additional geospatial modalities. The dataset includes aligned pairs of street view images, remote sensing images from Sentinel-2, cartographic data, and geographical coordinates. It was created by author DominikM198 and last updated on HuggingFace in May 2026.
Nemotron RL Instruction Following Structured Outputs V2 is a dataset for evaluating large language models on structured output generation. It was created by NVIDIA and last updated on June 4, 2026. The dataset includes two splits testing capabilities like freeform text generation and diversified tasks across multiple data formats.