Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
PKU-SafeRLHF-30K is a preference dataset containing over 30,000 expert comparison entries for AI safety research, created by the PKU-Alignment Team. Each entry includes a question and two responses with safety annotations. The dataset was last updated on November 20, 2023.
LLaVA-Plus-v1-117K is a set of 117,000 GPT-generated multimodal tool-augmented instruction-following data points. It was collected in September 2023 by prompting the ChatGPT/GPT-4-0314 API to build large multimodal agents with vision and language capabilities. The dataset was created by the LLaVA-VL organization.
CIVQA TesseractOCR LayoutLM is a dataset for Visual Question Answering on Czech invoices, created by fimu-docproc-research. The dataset was generated using Tesseract OCR and is pre-encoded for the LayoutLM model, focusing on 15 key invoice entities. It was last updated on November 21, 2023.
LSVQ is the largest dataset available for Non-reference Video Quality Assessment (NR-VQA), as stated in the description. This unofficial copy facilitates research after reports that original links are unavailable. The dataset was created by Ying et al. and published at CVPR in 2021.
A dataset for visual question answering on documents, published by HuggingFaceM4 on December 18, 2023. The dataset likely contains images of documents paired with questions and answers. Its specific scale, columns, and content require verification after download.
StackoverflowVQA-filtered-small is a multimodal dataset likely containing images and text for visual question answering tasks. The dataset was uploaded by mirzaei2114 to the Hugging Face platform and was last updated on December 2, 2023. Its specific size, contents, and license details are not provided in the available metadata.
Instruction Following Eval is a dataset for evaluating language models, published on Hugging Face by author wis-k. The dataset was last updated on December 5, 2023. Its specific content, scale, and structure require verification after download due to minimal metadata.
PKU-Alignment processed the HH-RLHF dataset into an easy-to-use conversational and human-preference form. The dataset was last updated on November 24, 2023. Its specific scale and column structure are not detailed in the provided metadata.
StackoverflowVQA is a multimodal dataset likely containing visual question-answering data. The dataset was uploaded by mirzaei2114 and was last updated on Hugging Face on November 29, 2023. The specific content, size, and structure are not detailed in the available metadata.
M3It provides between 1 million and 10 million bi-lingual instruction records for vision-language models, released by MMInstruction in 2023. It covers image classification and image-to-text tasks in both English and Chinese.
Cc2Dataset enables the extraction of multimodal pairs including image-text, audio-text, and video-text from the Common Crawl web archive. Developed by rom1504 and updated in 2023, it provides a pipeline to convert raw web documents into structured caption-media datasets. The tool is designed for big-data applications where media is paired with its surrounding document context.
A filtered dataset likely containing visual question-answering data derived from Stack Overflow content. It was published on the Hugging Face platform by the author mirzaei2114 and was last updated on December 2, 2023. The specific content, scale, and filtering criteria are not detailed in the available metadata.
A dataset used to train the CoEdIT text editing models, as described in the paper 'CoEdIT: Text Editing by Task-Specific Instruction Tuning'. It was created by authors Vipul Raheja, Dhruv Kumar, Ryan Koo, and Dongyeop Kang and is hosted on Hugging Face by Grammarly. The dataset was last updated on October 21, 2023.
96 challenging questions based on images from OpenImages form this evaluation benchmark for hallucination in Large Multimodal Models. It includes ground-truth answers and image contents. The dataset was created by Shengcao1006 and uploaded in November 2023.
A Korean language dataset constructed for supervised fine-tuning (SFT) of large language models as part of a Sungkyunkwan University industry-academic cooperation project. The dataset was created by preprocessing and filtering data from sources including Stanford Alpaca and OIG-Chip2 using ChatGPT-3.5 Turbo 16k to improve naturalness. The dataset page was last updated on 2023-09-25.
Image-caption pairs for logo designs were scraped from the logobook.com archive. The dataset was created by the author 'mozci' for a research project to fine-tune text-to-image diffusion models. The data was last updated on September 26, 2023.
OBELICS is a massive, curated collection of 141 million English web documents containing 115 billion text tokens and 353 million images. The documents feature interleaved text paragraphs and images, extracted from Common Crawl dumps. It was created by HuggingFaceM4 and released in August 2023.
Over 40 million images sourced from Wikimedia Commons comprise this collection curated by ryanrudes. Updated in October 2023, the repository provides a massive scale of visual data for deep learning and computer vision research.
A curated collection of human preference datasets across three categories: fine-tuning, RLHF, and evaluation. This repository indexes resources specifically designed for training and benchmarking Large Language Models against human-labeled preferences.
HuggingFaceM4 released MMBench_dev on August 23, 2023. It is a benchmark dataset designed to evaluate the performance of vision-language models, addressing challenges in assessing models like MiniGPT-4 and LLaVA. The dataset aims to move beyond traditional benchmarks such as VQAv2 and COCO Caption.