Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
The dataset connects 20,000 videos to temporally annotated sentence descriptions. On average, each video contains 3.65 temporally localized sentences describing unique segments and multiple events.
The Tumblr GIF (TGIF) dataset contains 100,000 animated GIFs and 120,000 descriptive sentences. GIFs were collected from randomly selected Tumblr posts published between May and June 2015, with sentences gathered via a crowdsourced annotation interface. It is designed for evaluating animated GIF and video description techniques.
The Tumblr GIF (TGIF) dataset contains 100,000 animated GIFs and 120,000 descriptive sentences. GIFs were collected from randomly selected Tumblr posts published between May and June 2015, with sentences gathered via a crowdsourced annotation interface. It is designed for evaluating animated GIF and video description techniques.
Aggregating captioned cartoons, combining image and text modalities. It was authored by juliaturc and last updated on November 8, 2022. The dataset is tagged with an US region focus and includes Parquet file formats.
Featuring 30,000 sarcastic tweets paired with GIF reactions. It was created for research on predicting induced affect, as detailed in an ACL 2021 paper by Shmueli, Ray, and Ku.
152,545 multiple-choice questions based on 21,793 video clips from 6 popular TV shows including The Big Bang Theory and Grey's Anatomy. The dataset provides paired subtitles and localized temporal annotations for every question to support multimodal reasoning.
TVQA+ provides spatio-temporal grounding labels for video question answering tasks. Developed by researchers for ACL 2020, the dataset facilitates multi-modal reasoning by linking natural language questions to specific video frames and regions.
Doodles Captions Blip is a dataset hosted on HuggingFace by julianmoraes, last updated in October 2022. The platform tags indicate it contains both image and text modalities, suggesting it likely contains pairs of doodle-style images and descriptive captions. The dataset's specific size, structure, and content require verification after download.
COYO-700M is a large-scale dataset containing 747 million image-text pairs with additional meta-attributes. It was created by KakaoBrain using a strategy of collecting informative alt-text and associated images from HTML documents. The dataset was last updated on August 30, 2022.
RedCaps is a dataset of 12 million image-text pairs collected from Reddit. It is designed for image-to-text tasks and was created by Karan Desai and colleagues.
Aggregating Flickr30K image caption quintets used to compute denotational similarities for semantic inference tasks. It was created by embedding-data and last updated in August 2022.
Title and encoded image pairs from Medium articles, derived from a Kaggle dataset of 128,000 articles. The images were centrally cropped to a square and resized to 256x256 pixels before being encoded into image tokens.
The Public Multimodal Dataset (PMD) contains 70 million image-text pairs with 68 million unique images. It was introduced in the FLAVA paper and aggregated from publicly-available sources including Conceptual Captions, WIT, Localized Narratives, RedCaps, COCO, SBU Captions, Visual Genome, and a subset of YFCC100M.
1,000 3D object models featuring synchronized visual, acoustic, and tactile data. The collection includes 3D meshes, simulated impact sounds, and high-resolution tactile images generated via the DIGIT sensor simulation.
The ActivityNet Captions dataset contains 20,000 videos, each annotated with an average of 3.65 temporally localized descriptive sentences, resulting in 100,000 total sentences. Each sentence describes a unique video segment and has an average length of 13.48 words. The dataset was created by Leyo.
Designed for text classification tasks, specifically sentiment classification, with a size category of 1K to 10K instances. It contains monolingual Russian text data, created by the author Aniemore.
DocVQA 1200 Examples is a multimodal dataset for visual question answering on documents. It contains 1,200 examples of images paired with text and questions, created by author nielsr and last updated in August 2022.
OK-VQA contains 14,055 open-ended visual questions, each with 5 ground truth answers. The dataset is manually filtered to ensure all questions require outside knowledge, such as from Wikipedia, and has been processed to reduce bias from common answers.
A collection of multimodal image-text pairs, with each sample including an image or image URL, associated text strings, a source identifier, and JSON-formatted metadata. The dataset was created by HuggingFaceM4 and was last updated in June 2022.
Offering image features extracted from the Flickr8k dataset using a ResNeXt-152 C4 architecture. It includes Arabic and English captions and splits provided by ElJundi et al., intended for use with the OSCAR learning method.