Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
The VQA dataset contains open-ended questions about images, requiring an understanding of vision, language, and commonsense knowledge to answer. It was created by HuggingFaceM4 and last updated in June 2022.
A dataset titled 'Fcd Lmv2' was authored by 'sheikh' and last updated on July 7, 2022. The dataset is associated with the tag 'Regionus', but no further descriptive details, column information, or row counts are available.
DocVQA Train is a dataset for visual question answering on document images. It was uploaded by Raagul04 to Hugging Face in July 2022. The dataset is intended for training models to answer questions based on visual content within documents.
ImageCoDe is a vision-and-language benchmark requiring contextual understanding of pragmatics, temporality, long descriptions, and visual nuances. The dataset was created by BennoKrojer and last updated on May 13, 2022. The specific row count, column count, and dataset size are unknown.
HowTo100M contains 136 million narrated video clips sourced from 1.2 million YouTube instructional videos spanning 15 years. The dataset focuses on videos where creators teach complex tasks, covering 23,000 activities in domains like cooking, crafting, and fitness.
This is version 1.0 of the ADVQA dataset, authored by HuggingFaceM4 and last updated in June 2022. The dataset's row count, column structure, and specific content are unknown.
The dataset comprises spatialized versions of the Libri-Trans and SLURP audio datasets, intended for enhancing translation and understanding tasks. It was authored by espnet and last updated in June 2022.
Over 10,000 artistic images from the WikiArt repository have been paired with descriptive captions generated by the BLIP model. This multimodal dataset was created by ChristophSchuhmann and uploaded to Hugging Face in May 2022. It combines visual art with machine-generated text descriptions.
4,000,000 image-caption pairs stored in PyArrow IPC format for high-performance multimodal training. The dataset utilizes memory-mapped files to enable low-latency data access during large-scale model optimization.
Multimodal Sarcasm Detection is a dataset for detecting sarcasm from multiple data modalities, likely combining text and visual information. The dataset was created by author Carol99 and was last updated on Hugging Face in April 2022. Specific details on the number of samples, features, and collection method are not provided in the available metadata.
25,000,000 image-caption pairs structured for large-scale multimodal model training. The collection expands upon the 4M Img Caps framework to provide a higher volume of text-image associations for vision-language tasks.
Filtered WIT is an image-text dataset derived from the Wikipedia Image Text (WIT) dataset, containing 10,000 samples per archived tar file. Each sample includes a .jpg image, a .txt caption, and a .json metadata file. The dataset is provided by LAION and was last updated in January 2022.
1,500,000 images representing Wikipedia entities curated for the Visual Question Answering over Entities (ViQuAE) benchmark. These images serve as a visual knowledge base for tasks requiring models to link visual inputs to external structured information and natural language questions.
3,700 question-answer pairs linked to images and a knowledge base of 1.5 million Wikipedia entities. The dataset facilitates visual entity retrieval where answers are specific entities rather than generic object labels.
DocVQA is a dataset for document visual question answering, created by nlpconnect and hosted on Hugging Face. The dataset was last updated in May 2022, though specific details on its size and composition are not provided in the available metadata.
3,700 question-answer pairs paired with images and a retrieval corpus of 1.5 million Wikipedia passages. The dataset focuses on entity-centric visual question answering, requiring models to identify visual entities and retrieve external knowledge to provide answers.
This collection aggregates multiple multimodal datasets and pre-computed visual features specifically curated for Visual Question Answering (VQA) and image captioning tasks. It provides a standardized interface for PyTorch users to access vision-language benchmarks through a dedicated Python package.
22 million compositional questions and 113,000 images featuring scene graphs. Structured semantic representations for both images and questions support multi-step visual reasoning and logic-based evaluation.
Built from open-ended questions paired with images, categorized by their requirement for vision, language, and commonsense reasoning. It provides a framework for testing multimodal understanding through tasks that cannot be solved by a single modality alone.
22 million compositional questions and 113,000 images featuring dense scene graph annotations. The dataset structures visual reasoning through functional programs that map out the logic required to reach an answer for each image.