Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,929 datasets
BALLADEER integrates EEG, eye tracking, and physiological signals from children and adolescents with ADHD and neurotypical controls. Its controlled protocol uses gamified cognitive tasks like Attention Slackline and CogniFit to elicit responses in attentional control and cognitive flexibility. This dataset supports the development of machine learning models for ADHD classification and the research of digital biomarkers.
LAD-Bench is a benchmark of more than 1,000 curated synthetic images designed to test the logical reasoning capabilities of Vision Language Models. It was created by SahasraK and introduced to address gaps in evaluating physical and social common sense for open-world AI deployment. The dataset was last updated on June 16, 2026.
Indian multilingual document images and OCR transcriptions curated by MILA: MULTILINGUAL INDIC LANGUAGE ARCHIVE. This representative subset contains samples spanning 19 Indian languages and scripts, focusing on real-world documents with complex layouts and noisy scans. The full dataset, covering all 22 official languages, is scheduled for release upon paper acceptance.
AnyAudio-Judge Bench is a bilingual (English/Chinese) multi-domain benchmark for evaluating instruction-audio alignment, released with the paper "AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following". It contains 7,920 curated samples per language across 7 subsets. The dataset was created by author cucl2 and was last updated on June 2, 2026.
Behavioral data on a large flock of flamingos collected by animal care staff after a change in their enclosure. The dataset includes a blank template for others to use. It is a 17.9 KB XLSX file authored by Paul Rose and last updated on 2026-05-19.
Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026). The dataset hosts 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Authored by kesav2k04, it was last updated on June 14, 2026.
1,119 sports video clips are paired with English captions authored and reviewed by expert labelers from Akai Space Labs. The dataset is designed for training and evaluating multimodal models. It was last updated on June 14, 2026.
A multi-modal dataset collected using the da Vinci Research Kit (dVRK). The dataset was created by jackzhy and a subset has been incorporated into NVIDIA's PhysicalAI-Robotics-Open-H-Embodiment collection. It was last updated on June 3, 2026.
Two spectral scenes, as depicted in the paper 'Bio-inspired multimodal imaging in reduced visibility' by PierreโJean Lapray. The dataset likely contains multimodal image data designed for research in computer vision under challenging visibility conditions. The specific data format, size, and collection details are not provided in the available metadata.
A multimodal dataset focused on driving behavior and human factors, likely containing eye-tracking data. It was authored by Xiaoming Tao and is available via the paperswithcode platform under an Open Access (green) license. The specific scale, collection period, and detailed contents are not provided in the available metadata.
Reaction time data from 20 participants and eye-tracking data from a subset of 10 participants, collected during an experimental study on alarm modalities for UAV signal-loss events. The dataset was contributed by Saleh, Nermeen to Harvard Dataverse and last updated in June 2026. It includes participant responses, alarm condition information, and eye-tracking metrics to support research on multimodal alarm systems.
A randomized controlled trial assessed the effect of a multimodal workshop on fifth-year medical students' clinical exam performance and anxiety. The study compared an intervention group receiving stress management and communication training against a control group, with anxiety measured using the STAI-State scale at multiple time points. The dataset likely contains OSCE scores and anxiety metrics for analysis.
12,483 images of analog clocks with time labels support training and evaluation of Vision-Language Models. The dataset originates from research presented at CVPR 2026 Findings. It was created by jaeha-choi and last updated on May 13, 2026.
Jiazhe Ma's dataset contains raw experimental data supporting a published article on cholesteric liquid crystal elastomer hollow fibers. The 72.0 MB dataset is organized by figure number from the manuscript, with each dataset presented as an Excel file or image. It was last updated on 2026-05-20 and is available under a CC-BY-4.0 license.
400,000 samples across five tasks were used to transfer a passive image editor into an autonomous, question-conditioned visual reasoning assistant. The dataset was created by BeichenZhang and last updated on 2026-05-25. It includes tasks such as Fine-grained Perception, Chart Understanding, Maze Solving, and Jigsaw Puzzle.
A multimodal dataset derived from the LLaVA-Instruct-150K source, containing synthetic annotations for tasks involving text, images, and speech. It is licensed under CC-BY-4.0 and was uploaded by author dreyn74. The dataset's size is indicated to be between 10,000 and 100,000 samples.
A dataset of approximately 20,000 rows containing instruction-following examples for language models. It is derived from the NVIDIA Nemotron-RL-Instruction-Following-Structured-Outputs-v2 dataset, with added thinking traces and validated final outputs. The dataset was created by electroglyph and last updated on June 17, 2026.
CapRL-Video-178K is a dataset providing file paths to over 97,000 video clips. The dataset is hosted by internlm on Hugging Face and was last updated on 2026-05-25. It serves as an index for video data from the LLaVA-Video-178K collection, which includes clips from sources like YouTube and ActivityNet.
OmniCap-IF is a benchmark dataset created by NJU-LINK for evaluating instruction following in omni-modal video captioning. It contains 480 videos and 1,920 instruction samples spanning tasks like understanding, generation, retrieval, and communication. Each sample pairs a prompt with fine-grained format and content checklists for evaluating structural, temporal, visual, audio, and audio-visual constraints.
Ayn-VQA is a culturally grounded Arabic multimodal evaluation dataset designed for the ImageEval 2026 Shared Task at ArabicNLP 2026. It tests whether a model can read a culturally specific image from a spoken Arabic question and distinguish grounded descriptions from plausible hallucinations. The dataset is authored by QCRI and was last updated in June 2026.