Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Conceptual 12M contains 12 million image-text pairs intended for vision-and-language pre-training. It was created by Google Research using a relaxed version of the data collection pipeline from Conceptual Captions 3M.
Conceptual Captions 12M (CC12M) contains 12 million image-text pairs designed for vision-and-language pre-training. It was created by pixparse and is a relaxed version of the CC3M dataset pipeline. The dataset instance was last updated on Hugging Face in December 2023.
The TextVQA dataset contains 45,336 questions based on 28,408 images from the OpenImages collection. It requires models to read and reason about text present within images to answer the provided questions.
A subset of approximately 15 million image-text pairs from the YFCC100M dataset, curated for training vision-language models. It was prepared by author vishaal27 and uploaded to Hugging Face in January 2024. The dataset provides page URLs and direct image download URLs for each entry.
464 multimodal earnings conference calls from S&P 500 companies featuring sentence-level alignment between audio recordings and text transcripts. The dataset provides structured financial disclosures paired with stock volatility labels for modeling market risk responses.
The AMTTL dataset is a monolingual Chinese text collection for token classification tasks, created by author gavinxing and last updated in January 2024. It is categorized as containing between 1,000 and 10,000 instances (1K<n<10K) and has crowdsourced annotations.
CogVLM-SFT-311K is the primary aligned corpus used in the initial training of CogVLM v1.0. The dataset contains approximately 311,000 bilingual visual instruction samples, constructed by selecting 3500 high-quality samples from MiniGPT-4, integrating them with LLaVA-Instruct-150K, and translating them into Chinese via a language model. The dataset was created by zai-org and last updated on December 26, 2023.
Mp DocVQA is a multimodal dataset for document visual question answering, created by the lmms-lab. It contains image-text pairs where questions are posed about document images. The dataset was last updated on Hugging Face in February 2024.
Conceptual Captions (CC3M) contains approximately 3.3 million images annotated with captions. The dataset was created by pixparse, with images and their raw descriptions harvested from the web, specifically from the Alt-text HTML attribute.
A dataset from Anthropic, published on HuggingFace by user 'nz' and last updated on February 2, 2024. The title suggests it contains data for Reinforcement Learning from Human Feedback (RLHF), a technique for aligning language models. The specific content, scale, and structure require verification after download.
Over 100,000 entries combine images with question-answer pairs for visual question answering tasks. The dataset was created by lmms-lab and last updated in January 2024.
12,000,000 English image-caption pairs derived from Google's Conceptual 12M dataset. The collection is structured in a TSV format containing image URLs, local filenames, and descriptive captions for each entry.
ViP-Bench is a region-level multimodal model evaluation benchmark curated by the University of Wisconsin-Madison. It provides two kinds of visual prompts for testing model understanding: bounding boxes and human-drawn diverse visual prompts. The dataset was last updated on December 15, 2023.
KREAM Product Blip Captions is a dataset for finetuning text-to-image generative models. It consists of image and text pairs collected from KREAM, a major online resale market in Korea. The dataset was created by author hahminlew and was last updated on December 7, 2023.
ChiMed-VL-Alignment is a multimodal dataset containing 580,014 Chinese medical image-text pairs. The dataset was created by author 'williamliu' and was last updated on the Hugging Face platform in December 2023. Pairs are categorized into context information and image-specific descriptions, with the context category containing 167 million tokens and descriptions containing 63 million tokens.
Latex Vlm is a dataset published on HuggingFace by JosselinSom. The dataset was last updated on January 20, 2024. Its specific content and scale are not detailed in the available metadata.
LanguageBind published a dataset titled 'Video Llava' on the HuggingFace platform in January 2024. The dataset likely contains video and text data for training or evaluating multimodal AI models. Specific details on size, format, and content are not provided in the available metadata.
Sam Llava Captions10M is a dataset published on HuggingFace by PixArt-alpha on January 12, 2024. The title suggests it contains image-caption pairs, likely for vision-language model training. The dataset's scale and specific content require verification after download.
OpenViVQA provides over 11,000 images paired with more than 37,000 open-ended question-answer pairs in Vietnamese. The dataset was created by uitnlp for the VLSP 2023 - ViVRC shared task challenge and was last updated in December 2023.
Human preference data collected from the r/WritingPrompts subreddit. The dataset was created by author euclaise and was last updated on December 25, 2023. The specific size, format, and column structure are not detailed in the provided metadata.