Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
A dataset published by lmms-lab on huggingface on 2024-06-28. The title suggests it relates to the LLaVA model and the Conceptual Captions 3M (CC3M) dataset, likely containing multimodal data for vision-language tasks. Specific content, size, and structure are not detailed in the provided metadata.
A dataset curated from Investopedia using a technique that scrapes unstructured data and employs an LLM to generate structured question-answer pairs. The dataset generation includes a self-verification method intended to reduce the probability of LLM hallucinations. The dataset was created by FinLang and was last updated on 2024-05-06.
OmniMedVQA is a large-scale Visual Question Answering benchmark for the medical domain, containing 118,010 images and 127,995 QA-items. The benchmark was collected from 73 different medical datasets and introduced by an author named foreverbeliever. It was last updated on 2024-04-30.
Over 700 anonymized images, primarily captured from vehicles, form this multimodal benchmark. Each image is paired with a question and a verifiable answer, designed to test real-world scene understanding. The dataset was released by xAI in April 2024.
A custom collection of paintings, images, and photographs exhibiting various types of damage. The dataset was created via manual collection and semi-automated annotation, with an initial sweep using the BLIP model followed by manual refinement. It was last updated on May 5, 2024, by the author 'calm-and-collected'.
LLM-jp provides a Japanese instruction-tuning dataset containing 33,000 entries. The dataset is a Japanese translation of a subset from the English OASST2 dataset, processed using DeepL. It was created by the LLM-jp collaborative project and last updated on April 28, 2024.
GuanacoDataset is a multimodal visual question answering dataset intended for aligning vision-language models with large language models. The dataset's creator is JosephusCheung, and it was last updated in April 2024, though its current availability on Hugging Face is uncertain.
ScreenSpot provides over 1200 text instructions paired with screens from iOS, Android, macOS, Windows, and web environments for evaluating GUI grounding. Researchers from Nanjing University and the Shanghai AI Laboratory created this benchmark to test large multimodal models. The dataset was last updated in April 2024.
Scientific Openly-Licensed Publications (SciOL) and its companion dataset, MuLMS-Img, are introduced in a WACV 2024 paper by Tim Tarsi et al. The dataset is designed for image-text tasks within the scientific domain and is hosted on HuggingFace by the author Timbrt. The dataset page was last updated on April 17, 2024.
DocVQA consists of 10,000 to 100,000 document images paired with question-answer sets, formatted by lmms-lab in 2024. This version is derived from the original 2020 DocVQA research to facilitate standardized evaluation of Large Multi-modality Models (LMMs). It provides a structured framework for testing how models interpret text and layout within diverse document types.
558,000 image-text pairs form this dataset for vision-language instruction tuning, curated by the lmms-lab research group. It was last updated in May 2024 and is hosted on Hugging Face. The data is specifically designed for training and evaluating multimodal AI models that process both visual and textual information.
ScreenSpot is an evaluation benchmark for GUI grounding created by researchers at Nanjing University and Shanghai AI Laboratory. It comprises over 1200 instructions from iOS, Android, macOS, Windows, and Web environments, along with annotations. The dataset was last updated on 2024-04-10.
A dataset named 'Small Clean Llava Instruct Mix' was published by the author 'damerajee' on the Hugging Face platform on 2024-05-27. The title suggests it is a curated collection of instruction-following examples, likely for training or fine-tuning vision-language models. Its specific content, size, and structure require verification after download.
TVR provides video-subtitle pairs and natural language queries for temporal moment retrieval, introduced by Jie Lei at ECCV 2020. The collection focuses on the TV show domain, requiring models to utilize both visual and textual dialogue features to locate specific events.
Screen2Words provides image captions for mobile application screens. It is built upon the RICO mobile app image database. The dataset was uploaded by rootsautomation to Hugging Face in April 2024.
Three categories of preference data—toxid-dpo-natural-v4, rawrr v2-1 stage 2, and no_robots—comprise this merged dataset. The samples focus on human-like conversational responses to prevent models from overfitting to rigid instruction-following templates.
TRL's Sentiment and Descriptiveness Preference Dataset originates from an early RLHF paper by OpenAI. The data has been preprocessed into a standard prompt, chosen, rejected format for reinforcement learning from human feedback. The dataset was last updated on the Hugging Face platform on 2024-04-09.
UltraInteract SFT is a large-scale, high-quality alignment dataset designed for complex reasoning tasks. The dataset, created by openbmb, includes preference trees with reasoning chains, multi-turn interaction trajectories, and pairwise data for preference learning. It was last updated on April 5,我们发现了一个问题,在生成 summary 时,我使用了
A multimodal dataset likely containing landscape imagery paired with compositional descriptions or labels. The dataset was authored by TomEijkelenkamp and published on the HuggingFace platform on May 22, 2024. Specific details regarding content, size, and structure are not provided.
A benchmark dataset for visual question answering tasks focused on table data. The dataset was uploaded by author 'terryoo' to Hugging Face and last updated on April 25, 2024. Its specific size, contents, and collection methodology are not detailed in the provided metadata.