Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,925 datasets
WorldEngineAI's WEB-Dataset is a large-scale, language-annotated real-robot bimanual manipulation dataset intended for post-training robotics foundation models. It spans 90 everyday manipulation tasks collected with a bimanual YAM follower arm teleoperated by a GELLO leader. The dataset records joint state, action, and three synchronized camera streams at 60 Hz.
A curated collection of 1,000 image and caption pairs. Each sample pairs an image with a detailed natural language description, making it suitable for training and evaluating vision-language models. The dataset was created by prithivMLmods and was last updated on July 4, 2026.
A replication package from an experimental study evaluating a multimodal chatbot as a pedagogical mediator for digital literacy among elderly women. The dataset includes anonymized participant data, task completion time records, success rates, and axial networks built from transcripts. It was authored by AMANDA SALES and last updated on 2026-05-30.
SovNodeAI's Certified Document QA dataset contains over 6,000 rows of machine-checkable question-answer claims for verifying large language model outputs. Every claim includes a certificate allowing item-by-item re-verification, and the data includes filings newer than major model training cutoffs. The dataset also includes a free 127,000-token verified long-context task set and a frontier failure table comparing six models on 100 questions.
187 image pairs from the AmalgaMatch dataset, partitioned into six distinct matching tasks and 19 material subsets, facilitate evaluation of foundation models for multimodal image registration. Ali Riza Durmaz published this supplementary PDF in May 2026 under a CC-BY-4.0 license. The dataset covers metals, alloys, and ceramics imaged with diverse microscopy modalities, presenting challenges like limited mutual information and field-of-view ratios as low as 2%.
Multimodal USElecDeb60To16 provides audio features and synthetic speech for U.S. presidential debates from 1960 to 2016. The dataset was created to augment pre-trained language models for argumentation mining, as described in a 2023 EACL Findings paper. It offers a version with large audio files and a lighter version without them.
Supplementary file 1 from a retrospective cohort study by Hui Zhang, published on figshare in 2026. The data likely contains results from 82 high-risk parturients receiving a multimodal analgesic protocol and 79 historical controls, collected between January 2023 and December 2024. Outcomes include postpartum depression incidence, Edinburgh Postnatal Depression Scale scores, Pittsburgh Sleep Quality Index scores, and opioid consumption.
Paired uterine whole-slide images and corresponding pathology reports for multimodal computational pathology research. The dataset was created by Zhengyang-TUM and is associated with a paper published on arXiv. The dataset page was last updated on July 17, 2026.
Catalogue metadata extracted from digitised material with a multimodal model and human-reviewed. The dataset is authored by the NationalLibraryOfScotland and was last updated on July 16, 2026. It contains structured fields describing index cards, including headings, types, and cross-references.
A curated subset of 35,794 image-caption pairs from the Conceptual Captions dataset, re-annotated in Russian for accessibility. The data was processed through semantic clustering of 2,484 groups and re-annotated using teacher vision-language models. It was created by Pavel Mikheyev and last updated in May 2026.
A multimodal dataset was used to develop a predictive model for overt hepatic encephalopathy (OHE) within one year after a transjugular intrahepatic portosystemic shunt (TIPS) procedure. The study by Lin-Feng Zhou, last updated in May 2026, integrated manual CT imaging features, radiomics, and clinical data from 338 patients treated between November 2015 and January 2022. The combined model (Model MRC) demonstrated superior predictive performance with an AUC of 0.902.
Lin-Feng Zhou's dataset supports a study developing a multimodal model to predict overt hepatic encephalopathy (OHE) within one year after a transjugular intrahepatic portosystemic shunt (TIPS) procedure. The data includes manual CT features, radiomics, and clinical data from 338 patients treated between November 2015 and January 2022. The combined model (Model MRC) achieved an area under the ROC curve of 0.902.
MODUS is a large-scale dataset with pixel-aligned samples across 15 modalities. The dataset includes modalities covering appearance, geometry, structure, segmentation, detection, text, and learned features. It was created by epfl-vilab-modus and was last updated on July 5, 2026.
A PDF supplementary file describes a decision-support framework for optimizing last-mile mail delivery in Australian regional areas. The study integrates mail demand and GIS data with an optimization engine to coordinate van, walking, and cycling routes. The system reportedly achieved reductions of up to 21.67% in delivery time and 11.36% in COโ emissions compared to van-only operations.
Indonesian-language Visual Question Answering dataset derived from VinDr-CXR radiologist annotations. It contains 15,991 questionโanswerโreason triples grounded in annotated findings. The dataset was created by Softcase and was last updated on 2026-07-09.
The SPHERE Challenge dataset was created for a 2016 machine learning competition held in conjunction with ECML-PKDD. It provides multimodal sensor data intended for human activity recognition tasks. The data was authored by Niall Twomey and colleagues from the SPHERE research project.
A Microsoft dataset release for the RESOURCE2SKILL system, which distills human-created multimodal resources into reusable executable skills for software agents. The dataset was last updated on 2026-07-17 and includes structured skill entries for discovery and inspection. The project page, paper, and code are available via the provided links.
A collection of datasets for training and evaluating machine learning models on small-molecule natural products. The data, totaling 128.0 MB, was compiled by Zhenming Liu from multiple public databases including COCONUT, NPASS, LOTUS, and MIBiG. The collection was last updated on 2026-04-30.
MMGist is a curated multimodal evaluation benchmark built from 18 widely used vision-language benchmarks. It contains 7,262 samples spanning seven capability dimensions and is designed to make LVLM evaluation more efficient, visually grounded, discriminative, and reliable. The dataset was authored by Winston-Yuan and last updated on June 29, 2026.
A multimodal dataset from a three-stage study examining color preference stability in spatial contexts. The data includes baseline preferences for ten Munsell hues, Preference and Comfort ratings, eye-tracking, and pupillometric data from a simulated makerspace environment, authored by Hourong Yu and last updated in May 2026. The dataset is shared under a CC-BY-4.0 license on figshare.