Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
5,040 text-image pairs across 13 safety scenarios including hate speech and illegal activities. The dataset provides a benchmark for evaluating the safety alignment of multimodal large language models. It specifically targets vulnerabilities in vision-language models through adversarial prompts.
DreamLIP-Long-Captions contains approximately 30 million image annotations consisting of detailed long captions. The captions were generated using pre-trained Multi-modality Large Language Models, with an average length of 247 characters.
MINT-1T is an open-source multimodal dataset containing 1 trillion text tokens and 3.4 billion interleaved images, representing a tenfold scale-up from prior open-source collections. It was created by a team from the University of Washington to support research in multimodal pretraining, incorporating sources like PDFs and ArXiv papers.
MINT-1T contains 1 trillion text tokens and 3.4 billion images, a tenfold scale increase from prior open-source multimodal collections. Created by a University of Washington team, this dataset interleaves text and images from sources including ArXiv papers and PDFs to support multimodal pretraining research.
A filtered version of the Dolly dataset, designed for instruction tuning of large language models. The dataset was created by qingy2024 and was last updated in November 2024. It contains text data categorized for fine-tuning tasks.
1.432 million image-QA instances developed by wentao-yuan in 2024 facilitate fine-tuning Vision-Language Models for spatial affordance prediction. The collection integrates 667K synthetic instances for object and free space referencing with 100K LVIS detection samples and 150K instruction-following pairs.
GenAI-Bench is a benchmark for evaluating multimodal large language models' ability to judge the quality of AI-generated content. The dataset is based on human preference data collected via the GenAI Arena platform and is maintained by TIGER-Lab. It was last updated on 2024-09-08.
Mimic Cxr Vqa likely contains chest X-ray images paired with questions and answers for visual question answering tasks. Published on huggingface by MiniMedMind on November 3, 2024, its exact size and content are unspecified.
Over 250 million human ratings on more than 2.2 million cartoon captions form a multimodal preference dataset for creative tasks. The dataset was created by researchers and is associated with a paper titled 'Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning'. It was last updated on the platform in September 2024.
AlignMMBench is a multimodal alignment benchmark released in June 2024 by zai-org. It evaluates Chinese large vision-language models across single-turn and multi-turn dialogue scenarios. The dataset encompasses three categories and thirteen sub-tasks, as detailed in its associated arXiv paper.
A modified version of the Amazon Multimodal Product dataset, slimmed for training multimodal LLMs. The dataset includes product descriptions generated using the Gemini Flash model. It was created by philschmid and last updated in September 2024.
1.6 million biomedical image-caption pairs were collected from the PubMedCentral OpenAccess subset to address data scarcity for foundation models. The dataset, released by author axiong in August 2024, is described as 8 times larger than previous collections. It features fine-grained alignment between subfigures and subcaptions, covering diverse modalities and diseases.
Lmsys Arena Human Preference 55K Sharegpt is a dataset published on HuggingFace by mlabonne and last updated on October 18, 2024. The title suggests it contains 55,000 records of human preference judgments, likely sourced from the LMSys Chatbot Arena or ShareGPT platforms. The dataset's specific content and structure require verification after download.
Annotations from the VAST-27M dataset created for the 2024 paper "VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset". The dataset was derived from work by researchers at the University of Chinese Academy of Sciences and the Institute of Automation, Chinese Academy of Science. It was uploaded to the Hugging Face platform by the user 'it-just-works' on September 10, 2024.
InfinityMATH is a scalable instruction tuning dataset for programmatic mathematical reasoning. The dataset was created by BAAI and was last updated on September 3, 2024. Its construction pipeline emphasizes decoupling numbers from problems to synthesize number-independent programs.
The dataset integrates images from the Infographic_vqa and AFTDB (Arxiv Figure Table Database) collections. It consists of image-text pairs, with each image linked to an average of five questions and answers available in both English and French. The dataset was created by cmarkea and last updated on Hugging Face in August 2024.
Llava Recap Cc12M is a multimodal dataset created by lmms-lab and published on Hugging Face on October 10, 2024. The title suggests it likely contains image-text pairs for instruction-following tasks. The dataset's specific content, size, and structure require verification after download.
Over 8,700 labels and descriptions were generated for 1,252 Vietnamese handwriting images using the Gemini 1.5 Flash model. The dataset was created by 5CD-AI from the train splits of the Cinnamon AI Challenge and UIT-HWDB datasets. It was last updated on Hugging Face in August 2024.
LLaVA-NeXT Data contains between 100,000 and 1,000,000 instruction-tuning pairs for multimodal large language models, released by lmms-lab in August 2024. It provides the specific data mixtures used to train the LLaVA-NeXT and LLaVA-NeXT (stronger) models, featuring synchronized image and text instruction sets.