Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
SMolInstruct is a large-scale instruction tuning dataset for chemistry tasks centered around small molecules. It contains over 3 million samples across 14 distinct chemistry tasks. The dataset was created by osunlp and last updated on Hugging Face in September 2024.
9,594 Vietnamese images from the ViTextVQA train split were analyzed using the Gemini 1.5 Flash model to generate over 50,000 detailed descriptions, questions, and answers. The dataset was created by 5CD-AI and was last updated on August 25, 2024. It is hosted on the Hugging Face platform.
UltraMedical-Preference is a dataset from TsinghuaC3I that enhances the UltraMedical collection with preference annotations. It includes responses sampled from both open-source and proprietary models, annotated for user preferences. The dataset was last updated on August 20, 2024.
BioMed-VITAL Instructions is a dataset for tuning multimodal AI models on biomedical visual tasks with clinician preference alignment. It contains multiple files ranging from 60,000 to 210,000 instruction samples, with file sizes from 127 MB to 463 MB. The dataset was created by authors including Hejie Cui, Lingjun Mao, and Carl Yang, and was last updated on August 17, 2024.
Over 6,000 detailed descriptions and query-based questions and answers generated for 1,056 Vietnamese images. The annotations were produced by the Gemini 1.5 Flash model using images from the VinText dataset train split. The dataset was created by 5CD-AI and last updated on August 25, 2024.
56,989 images depicting quintessentially Vietnamese scenes were annotated using Visual Question Answering (VQA) techniques. The dataset includes landscapes, historical sites, culinary specialties, festivals, and everyday life from various regions. It was created by 5CD-AI and last updated on Hugging Face in August 2024.
This multimodal fashion dataset provides image-text pairs annotated across categories, style, colors, materials, keywords, and fine-details. It is specifically curated to evaluate vision-language models like Marqo-FashionCLIP and Marqo-FashionSigLIP using fine-grained attribute metadata.
A collection of 4 million Hebrew image captions derived from the Recap-DataComp-1B dataset, published by NVIDIA on August 27, 2024. It is designed for training vision-language models like CLIP for the Hebrew language, providing text captions paired with references to pre-computed image embeddings.
27,519 images and corresponding question-answer pairs translated from the GQA train_balanced and testdev_balanced splits into Russian. The data underwent gpt-4-turbo translation followed by manual validation to correct errors and remove safety-filtered content. It is structured for use within the lmms-eval pipeline to support multimodal model benchmarking.
A Multi-Choice Visual Question Answering dataset designed to evaluate Vision-Language Models on their understanding of Korean culture. It was created through a Human-VLM collaboration and is part of research presented in a June 2024 arXiv paper. The dataset was last updated on HuggingFace on August 17, 2024.
300,000 examples of visual instruction data for training multimodal large language models. The dataset combines 150,000 English examples from the LLaVA project and 150,000 from the openbmb project. Author BUAADreamer uploaded this collection to Hugging Face on September 2, 2024.
Chatbot Arena Conversations JA (calm2) is a Japanese instruction dataset constructed for RLHF, as described in its associated paper. The dataset was created to test whether English datasets can be adapted for Japanese using only open-source tools and models. Prompts are Japanese translations of user inputs from the lmsys/chatbot_arena_conversations dataset, which are human-written and licensed under CC-BY 4.0.
A multimodal dataset of physics problems designed for chain-of-thought reasoning. It contains 2,100 problems across three domains: 1,000 on Kinematics, 600 on Electricity and Circuits, and 500 on Thermodynamics. The dataset was created by Vikhrmodels and last updated on August 4, 2024.
141 million interleaved image-text web documents containing 115 billion text tokens and 353 million images comprise the OBELICS collection. Created by Hugging Face and updated in 2024, it serves as a massive open-source resource for multimodal AI development.
A collection of instruction and toxic alignment datasets for 14 Indic languages, created by ai4bharat and last updated on July 25, 2024. The datasets include subsets like IndicAlign-Instruct, Indic-ShareLlama, and IndicAlign-Toxic, which were translated using IndicTrans2. The full curation process is detailed in an associated arXiv paper.
10,000,000 image-caption pairs generated using the Florence-2 vision-language model for the Megalith-10M image collection. Textual descriptions supplement the previously uncaptioned CC-0 like images to support vision-language model training.
Argilla's 7,000-pair dataset, built with the distilabel tool, is designed for Direct Preference Optimization (DPO) training of chat models. This preview version, released on July 16, 2024, is based on the LDJnr/Capybara dataset and aims to address the scarcity of multi-turn dialogue preference data used in major RLHF works. A full version with more model responses is planned for a future release.
Hindi-language dataset for visual question answering tasks, published on Hugging Face by author 'azharumo'. The dataset was last updated on September 17, 2024. Its specific size, structure, and annotation details require verification after download.
SPA-VL contains 100,788 samples across 6 harmfulness domains and 53 subcategories, released by researcher sqrti in mid-2024. The dataset facilitates safety preference alignment for Vision Language Models (VLMs) using multimodal image-text pairs.
A dataset designed for instruction tuning in multimodal settings involving visual interaction data. It was created by nyu-visionx and released in 2024 to address the scarcity of high-quality multimodal instruction-tuning data. The dataset aims to maintain the language abilities of multimodal large language models.