Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
IFBench provides a benchmark for evaluating reward models designed to assess instruction-following capabilities in AI agents. The dataset was created by the THU-KEG research group and was published in March 2025 alongside their paper on agentic reward modeling. It contains samples with unique identifiers and source annotations for structured evaluation.
Egotextvqa is a multimodal dataset for video question answering tasks. The dataset was created by ShengZhou97 and was last updated in April 2025. It contains video and text data, focusing on reasoning tasks that require understanding both visual and language information.
1-hour videos and v1.0 development set annotations for long-form video-language understanding. This benchmark from Stanford University was introduced at NeurIPS 2024 to evaluate models on extended temporal sequences.
DriveLM-Data comprises two distinct components: DriveLM-nuScenes and DriveLM-CARLA. The dataset is designed to facilitate Perception, Prediction, Planning, Behavior, and Motion tasks with human-written reasoning logic. It was created by OpenDriveLab and was last updated on March 4, 2025.
A Thai translation of the LLaVA-CC3M-Pretrain-595K dataset, originally created by Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. This dataset is intended for pre-training large multimodal models with Thai language capabilities and was uploaded to Hugging Face by user 'worapob' on March 9, —2025.
PuzzleVQA is a dataset created by declare-lab for evaluating large multimodal models. The dataset likely contains puzzles based on abstract patterns to test general intelligence and reasoning abilities. It was last updated on Hugging Face on February 26, 2025.
MMIR is a benchmark dataset designed to test multimodal large language models' ability to detect real-world cross-modal inconsistencies. It contains 534 carefully curated samples, each with a single semantic mismatch across five error categories, spanning webpages, slides, and posters. The dataset was created by rippleripple and last updated on 2025-02-25.
WMT24++ Images provides source URLs and full-page document screenshots for the translation data used in the WMT24++ project. The dataset, created by Google and last updated on 2025-02 24, preserves original document structure with embedded images. It is intended to support multimodal translation and language understanding research.
A multimodal dataset likely designed for instruction-tuning of Vision-Language Models (VLMs). The dataset was published on HuggingFace by Hirai-Labs and was last updated on April 4, 2025. Its specific content and scale are not detailed in the available metadata.
Comprising human preference data for text-to-video generation, collected via the Rapidata API in approximately 12 hours. It is used to benchmark five AI models: Sora, Hunyouan, Pika 2.0, Runway ML Alpha, and Luma Ray 2. Row and column counts are unknown.
PubMedVision is a large-scale medical visual question answering dataset built from image-text pairs extracted from PubMed. FreedomIntelligence enhanced the data quality using GPT-4V and added annotations for body parts and modality. The dataset was updated in February 2025.
Harveyaot published a dataset of HTML visual design preferences on the Hugging Face platform on April 8, 2025. The dataset likely contains 1,000 samples, as suggested by the title. Its specific content and structure require verification after download.
A collection of multimodal mathematics problems and reasoning chains presented in the URSA research paper. The dataset was created by the URSA-MATH organization and was last updated on February 18, 2025. It likely contains over one million examples integrating visual and textual data for training and evaluating AI models.
IndicMMVet is a dataset for evaluating Large Vision-Language Models on Visual Question Answering tasks, created by krutrim-ai-labs. It focuses on integrated capabilities and multilingual content, specifically for Indian contexts. The dataset was last updated on March 5, 2025.
A dataset published on huggingface by henry-07 on April 6, 2025. It likely contains pairs of satellite imagery from the Sentinel program and corresponding textual captions. The specific volume, format, and column structure are unknown.
73,893 short videos from the TRECVID VTT task, each ranging from 3 to 10 seconds in duration. The dataset includes between 2 and 5 human-written captions per video, created by dedicated annotators hired by NIST.
100,000 image conversation samples derived from 45,000 web documents in the OBELICS dataset. GPT-4V and OpenChat 3.5 were used to generate contextual captions and convert them into diverse free-form conversations. The dataset was authored by tiiuae and last updated on February 17, 2025.
VisCon-100K is a dataset of 100,000 image-conversation samples designed for fine-tuning vision-language models. It is derived from 45,000 web documents in the OBELICS dataset, with captions generated by GPT-4V and converted into free-form conversations by OpenChat 3.5. The dataset was created by tiiuae and last updated on February 17, 2025.
BLIP3-OCR-200M is a dataset designed to improve Vision-Language Models' ability to process text within images. It was created by Salesforce and was last updated on February 3, 2025. The dataset likely contains images integrated with Optical Character Recognition (OCR) data to address limitations in interpreting documents and charts.
Art Museums PD 440K is a dataset for training text-to-image and multimodal models, containing images and captions sourced from public domain or CC0-licensed materials. The dataset includes English captions translated to Japanese using the ElanMT model, which was trained on licensed corpus. The creator is Mitsua, with the dataset last updated on February 13, 2025.