Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
595,000 image-text pairs form a subset of the CC-3M dataset, filtered for balanced concept coverage. It was created by liuhaotian for the pretraining stage of visual instruction tuning, aiming to build large multimodal models. The dataset was last updated on July 6, 2023.
262,110 natural language captions describing 108,965 video segments from 6 popular TV shows. The dataset facilitates multimodal video captioning by providing visual frames alongside time-aligned subtitle dialogue.
A subset of the LAION/CC/SBU dataset filtered for more balanced concept coverage distribution, constructed for the pretraining stage of visual instruction tuning. It contains synthetic captions generated by BLIP for reference and aims to build large multimodal models towards GPT-4 vision/language capability. The dataset was created by liuhaotian and last updated in July 2023.
LLaVAR provides a collection of 422,000 pretraining and 16,000 to 20,000 instruction-following data pairs for training multimodal AI models. Created by SALT-NLP, this dataset enhances visual instruction tuning by focusing on images containing text. The dataset was released and last updated in July 2023.
4,241 multimodal science questions representing the test split of the ScienceQA benchmark. It contains image-based multiple-choice questions accompanied by hints, lectures, and step-by-step explanations across natural, social, and language science subjects.
Tasksource provides the OASST1 dataset preprocessed for reward modeling. It contains pairwise human feedback data for training reinforcement learning from human feedback (RLHF) reward models, focusing on conversational AI and multilingual text.
Psychology RLHF data was used to train a LLaMA-7B reward model. The dataset was uploaded by author 'samhog' to Hugging Face on July 17, 2023. Its specific content, size, and structure are not detailed in the provided metadata.
Anthropic's HH-RLHF dataset contains between 100,000 and 1,000,000 human preference comparisons focused on model helpfulness and harmlessness, released in 2022. These text-based records are designed to facilitate the training of reward models for Reinforcement Learning from Human Feedback (RLHF) rather than supervised fine-tuning.
ImageRewardDB is a text-to-image human preference dataset containing 137,000 expert comparison pairs. It was created by zai-org and uploaded to Hugging Face on June 21, 2023. The dataset is built from text prompts and corresponding model outputs sourced from DiffusionDB.
PathVQA is a dataset for Medical Visual Question Answering built from the 'Textbook of Pathology' and 'Basic Pathology' textbooks. It contains question-answer pairs on pathology images, including both open-ended and binary yes/no questions.
Clinician-generated question-answer pairs paired with radiology images across open-ended and binary 'yes/no' categories. The dataset utilizes medical imagery sourced from the MedPix open-access database to support the development of Medical Visual Question Answering (VQA) systems.
13,003 images of 11,003 identities accompanied by 80,440 natural language descriptions. The dataset facilitates cross-modal person search by linking visual pedestrian data from surveillance cameras with detailed textual attributes.
A dataset for instruction tuning with GPT-4, created by the team referenced in the citation. The dataset page was last updated on 2023-05-03. It is intended for research use only and is licensed under CC BY NC 4.0.
A test set for the OK-VQA (Outside Knowledge Visual Question Answering) benchmark, created by Multimodal-Fatima and uploaded to Hugging Face on 2023-05-29. The dataset is designed for evaluating models that answer questions about images using external world knowledge. Specific details on size, columns, and license are not provided in the metadata.
Japanese Hh Rlhf 49K is a dataset derived from kunishou/hh-rlhf-49k-ja, excluding examples where ng_translation equals 1. The dataset was authored by fn-aka-mur and last updated on Hugging Face in May 2023. Its specific size and row count are not detailed in the provided metadata.
VQAv2 is a dataset for visual question answering tasks, uploaded to Hugging Face by landersanmi. The dataset was last updated on June 2, 2023. Specific details on size, columns, and license are not provided in the available metadata.
Instruction Tuning with GPT-4 is the title of the associated research paper. The dataset was created by the team llm-wizard and last updated on April 7, 2023. It is licensed for non-commercial research use under CC BY NC 4.0.
VQAv2_train is a dataset for visual question answering tasks, likely containing pairs of images and questions with corresponding answers. The dataset was uploaded by Multimodal-Fatima to Hugging Face and last updated in April 2023.
Image-text pairs for Italian Contrastive LanguageโImage Pre-training (CLIP). This data aligns visual representations with Italian linguistic descriptions to support cross-modal retrieval and zero-shot classification.
Face Synthetics Spiga Captioned is a copy of the Microsoft FaceSynthetics dataset enhanced with SPIGA-calculated facial landmark annotations and BLIP-generated text captions. The dataset, created by multimodalart and last updated in March 2023, is designed for multimodal tasks involving synthetic facial imagery.