Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,923 datasets
A pelvic MRI dataset of 74 subjects and 3,449 T2-weighted slices from two institutions for developing AI models for uterus segmentation in endometriosis. The dataset was used to fine-tune the Endo-MedSAM model, achieving mean 3D Dice scores of 0.81β0.88 with bounding-box prompts. The dataset was uploaded by Rawan AlSaad on figshare in May 2026.
65 sessions across 422 segments with 4,849 files provide synchronized pose trajectories for EdgeXR and VR research. The dataset includes temporally and spatially aligned pose data captured at 500 Hz from SteamVR gaming sessions via OpenXR API readings and marker-based optical motion capture. Ziyu Zhong organized the data into a cleaner structure and prepared it for confidential peer review on Harvard Dataverse.
17 healthy participants (7 females, 10 males, aged 19β34) performed walking activities across diverse indoor and outdoor terrains. The dataset includes motion data from 7 inertial sensors, foot pressure from 96-point force sensors, and visual data from 3 front-facing cameras, all annotated with 16 locomotion state classes. Collected by Chen Wang and shared under a CC-BY-4.0 license, this 16.9 GB dataset was last updated on 2026-05-31.
Renjie Lu developed a multimodal model integrating tumor radiomics and lymph node morphology for predicting axillary nodal metastasis burden in breast cancer. The dataset includes information from 583 patients with pathologically confirmed breast cancer, split into training and testing cohorts. The model was last updated on June 4, 2026.
2,956 Ultra High Definition (UHD) image samples are paired with rich, long-form captions for vision-language research. The dataset, created by prithivMLmods, is designed for tasks like image understanding and dense captioning. It was last updated on July 15, 2026.
ABC-130k is a multimodal dataset of bimanual robot teleoperation episodes. It contains 134,806 episodes across 195 tasks, representing 3,553 hours of synchronized multi-camera video and robot telemetry. The dataset was created by Voxel51 and is hosted on Hugging Face.
A preclinical study by Hyo Jin Kim from Harvard Dataverse, last updated in 2026, evaluates a drug delivery system for pain management. The dataset likely contains results from 64 Sprague-Dawley rats across four treatment groups, measuring mechanical withdrawal thresholds and inflammatory cytokine levels over one week. It focuses on the analgesic and anti-inflammatory effects of combining ReproGel with ropivacaine 0.375%.
48 participants completed a within-subjects mixed reality experiment using a Meta Quest Pro headset. The dataset includes perceived intrusiveness, physiological arousal, embodied avoidance behaviors, and cognitive performance metrics, all timestamp-synchronized at the trial level. Authored by Yuxuan Li and hosted on Harvard Dataverse, it was last updated in July 2026.
10,000 images are paired with detailed, long-form captions generated by the Qwen3.5 multimodal model. The dataset is designed for dense image captioning, with descriptions focusing on scene composition, subject attributes, and spatial relationships. It was created by prithivMLmods and last updated on July 13, 2026.
Micro-OD is a benchmark of 252 images curated for in-context learning, with bounding-box annotations for 11 cell types across four sources. It was created by Shreyan Ganguly and last updated in May 2026. The dataset is designed to evaluate vision-language models for few-shot object detection in biomedical microscopy.
A video benchmark collected from blind individuals for evaluating AI assistance models. The dataset, created by MCG-NJU, is designed for tasks like Proactive Reminder and Visual Question Answering. It was last updated on 2026-07-17.
A 2026 evaluation assesses the capabilities of foundation models like MatchAnything RoMa and ELoFTR for multimodal image matching in materials science. The analysis uses the AmalgaMatch dataset, which contains 187 image pairs across six distinct matching tasks and 19 different materials. The work was authored by Ali Riza Durmaz and is shared under a CC-BY-4.0 license.
Crowdsourced typing preference data from a study that derived ergonomics objectives from user preferences. The dataset includes materials for the Engram approach to optimizing keyboard layouts for English and Spanish, created by Arno Klein and last updated in May 2026. It contains data, software, documentation, and layouts totaling 10.6 MB.
Vietnamese multimodal reasoning data featuring multi-turn visual question-answering grounded on natural images. Each answer includes an explicit chain-of-thought reasoning trace, synthesized by the Qwen3.5-397B-A17B model over images from the LAION-derived Vi-Laion-gemini-VQA set. The dataset was curated by TrαΊ§n Nhiα»m and last updated on 2026-07-17.
Vietnamese document-image understanding with explicit reasoning chains for multi-turn question-answering. The dataset is based on scanned or rendered Vietnamese document pages such as textbooks, articles, and worksheets. It was curated by TrαΊ§n Nhiα»m and the reasoning and answers were synthesized by the Qwen3.5-397B-A17B model over the Viet-Doc-VQA-II document collection.
VSI-Super-Wild is a benchmark for evaluating multimodal models on spatial supersensing capabilities in long-form, in-the-wild videos. It was created by researchers from Tsinghua University, NVIDIA, and Stanford University for the ECCV 2026 conference. The dataset moves beyond short indoor clips and object-centric settings to study world state maintenance and prediction.
Twenty individuals with mild traumatic brain injury and 24 healthy controls underwent advanced diffusion MRI and cognitive assessment. The data includes multi-shell DTI, free-water corrected DTI, diffusion kurtosis imaging, and NODDI metrics, linked to MoCA and GOS-E clinical scores. Authored by Maurizio Bergamino and shared under CC-BY-4.0, this dataset was last updated on May 28, 2026.
Eighty-seven subjects with at-risk mental states (ARMS) were followed up, with clinical outcomes classified into four ordered categories. The dataset contains baseline measures for 15 explanatory variables, including clinical symptoms, cognitive functioning, and electrophysiological measures like P300 and mismatch negativity. The data was authored by Kazuya Nagasawa and last updated on 2026-05-28.
AI-CVM's Cardiac-CT dataset accompanies a research paper on a unified framework for cardiac CT segmentation and phenotyping. The dataset was used for human-in-the-loop annotation, vision foundation model development, and multicenter evaluation. It was last updated on July 15, 2026.
WorldEngineAI's WEB-Dataset is a large-scale, language-annotated real-robot bimanual manipulation dataset intended for post-training robotics foundation models. It spans 90 everyday manipulation tasks collected with a bimanual YAM follower arm teleoperated by a GELLO leader. The dataset records joint state, action, and three synchronized camera streams at 60 Hz.