Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Image-text pairs from the MS COCO 2017 dataset, sourced from cocodataset.org. The data is provided in two formats: a dense format with several sentences per image row and a long format with one caption per row, expanding the dataset length by a factor of five.
Presenting a reformatted version of theblackcat102/llava-instruct-mix, prepared for Vision Supervised Fine-Tuning (VSFT) with the TRL SFT Trainer. It is designed for instruction tuning of multimodal vision-language models. The dataset's author is HuggingFaceH4, and it was last updated in April 2024.
Polaris contains between 100,000 and 1,000,000 records of human feedback on image-caption pairs, released by researcher yuwd in 2024. This multimodal dataset supports the development of evaluation metrics that align with human judgment as described in the CVPR 2024 paper "Polos."
CaptionEmporium provides 6.92 million captions for safe-for-work images from the e621/e926 platform, extending to January 2023. The dataset includes captions generated by a large language model (mistralai/Mistral-7B-v0.1) and a multimodal model (THUDM/CogVLM), with 8 LLM and 1 CogVLM caption per image. Most captions are described as substantially larger than 77 tokens.
Supplementary materials for a study on the effects of scale on multimodal deixis. The dataset includes gesture form coding from the study's first author and a reliability coder, along with annotation guidelines and an R script for statistical replication. The data was archived in the Texas Data Repository and last updated in March 2024.
24,903 visual question-answering pairs paired with images from the COCO dataset, categorized into multiple-choice and direct-answer formats. Each entry includes human-annotated rationales explaining the reasoning required to answer questions that necessitate external knowledge beyond the visual content.
Pokémon BLIP captions is a multimodal dataset used to train a Pokémon text-to-image model. The dataset was created by author reach-vb and last updated on March 12, 2024. It contains Pokémon images from the FastGAN project paired with captions generated by the pre-trained BLIP model.
Hugging Face released Chug in April 2024 to provide sharded dataset loaders and decoders for multi-modal document, image, and text data. It focuses on efficient distributed training using WebDataset and PDF formats for computer vision and document understanding tasks.
VizWiz-VQA is a large-scale dataset for evaluating large multi-modality models. It is a formatted version used in the lmms-eval pipeline for one-click model evaluations. The dataset was created by lmms-lab and was last updated on March 8, 2024.
Ai2D contains between 1,000 and 10,000 scientific diagrams with corresponding text annotations, published by Aniruddha Kembhavi and the Allen Institute for AI in 2016. The dataset is designed to support research in diagrammatic reasoning and visual question answering within the scientific domain.
A formatted version of the TextVQA benchmark dataset, used for evaluating large multi-modality models. It was created by lmms-lab and last updated on March 8, 2024. The dataset is part of the lmms-eval pipeline for one-click model evaluations.
A blend of publicly available datasets for instruction tuning, including samples from OASST, CodeContests, FLAN, T0, Open_Platypus, and GSM8K. The dataset was created by NVIDIA and last updated on March 9, 2024. It consists of four columns, though specific column names and the total number of rows are not detailed in the provided metadata.
A formatted evaluation suite for large multi-modality models (LMMs), created by lmms-lab and last updated on March 8, 2024. It is designed to accelerate LMM development by enabling one-click evaluations through the lmms-eval pipeline. The dataset is based on the MM-Vet benchmark described in the associated research paper.
A dataset of Pokemon image-text pairs was removed from the Hugging Face platform on March 20, 2024. The takedown was initiated by The Pokémon Company International, Inc. via a DMCA notice, and the dataset author is listed as 'lambda'.
MMInstruction created a dataset for multimodal question answering, likely pairing images of scientific figures from arXiv papers with multiple-choice questions. The dataset was last updated on March 5, 2024. Each example includes an image path and a set of answer options, suggesting a focus on visual reasoning in academic contexts.
This dataset supports the FETA research paper, which was published as a main conference paper at NeurIPS 2022. It is used for specializing foundation models for expert task applications, with the official resources available on a dedicated GitHub repository.
12.8 million image URLs and their corresponding CLIP embeddings derived from the datacomp_small benchmark. The dataset is processed via the Fondant framework to provide a production-ready format for multimodal machine learning tasks without requiring raw image storage.
A visual dataset of emoticons annotated using the image parsing capabilities of the glm-4v and step-1v multimodal AI models. The dataset was created by LLM-Red-Team and was last updated on April 27, 2024. The specific number of images, rows, and columns is unknown.
Parsa-ra developed this multi-modal dataset interface on GitHub, with the last update recorded in April 2024. It serves as a practice for unified data handling, though specific record counts and file formats are currently undocumented in the repository metadata.
28,408 images from Open Images paired with 142,040 captions that require models to read and reason about text within the visual scene. This version is specifically formatted for the lmms-eval pipeline to facilitate standardized benchmarking of large multi-modality models.