Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
2023 images were pulled from Pexels, primarily depicting people holding objects. The dataset includes full images paired with captions generated by the CogVLM model. It was created by author 'lodestones' and last updated on the platform in June 2024.
SLAKE is a Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering, presented at ISBI 2021. This version, uploaded by mdwiratathya, filters the original bilingual dataset to contain only English entries, providing images as PIL objects, questions, and answers. The dataset was last updated on the Hugging Face platform on June 14, 2024.
A collection of controversial and adult-themed images for training multimodal detection models, curated by QuixiAI. The dataset is designed to enable detailed categorization and filtering of such content.
Created by zjunlp for ICLR 2023, this dataset supports multimodal analogical reasoning over knowledge graphs. It provides a structured environment for the MARS framework, linking visual and textual data to relational graph structures for reasoning tasks.
Pexels provided over 10,000 photographs of buildings and unique architecture in 2023. The dataset creator 'lodestones' used the CogVLM model to generate descriptive captions for each image. The dataset was last updated on the Hugging Face platform in June 2024.
CaptionEmporium's anime-caption-danbooru-2021-sfw-5m-hq dataset contains 5.71 million captions for 1.43 million safe-for-work anime-style images from the Danbooru 2021 dataset. Each image has four captions generated by different AI models, including CogVLM and variants of LLaVA-v1.6-34b. The dataset was last updated on Hugging Face in June 2024.
This multimodal agent benchmark evaluates AI performance within simulated clinical environments using language agents. It adapts the MedQA dataset to facilitate interactive diagnostic reasoning between AI doctors and simulated patients across various medical scenarios.
A 2024 mixture of text preference datasets used to train the weqweasdas/RM-Mistral-7B reward model for Reinforcement Learning from Human Feedback. The dataset was created by OpenRLHF and includes multiple sources of human-annotated comparisons. It is designed for training models to score and rank text outputs based on human preferences.
Apple provides metadata for the TiC-DataComp benchmark, which evaluates time-continual learning for image-text models. The dataset contains timestamp groupings by year and month for DataComp-1B images, sourced from CommonCrawl. It also includes unique identifiers for TiC-DataCompNet and TiC-DataComp-Retrieval evaluations of CLIP models.
A multimodal dataset published on huggingface by MAINLAND on July 22, 2024. The title suggests it likely contains satellite imagery paired with textual instructions, intended for training vision-language models. The specific content, scale, and geographic scope require verification after download.
This audio-text dataset provides paired audio signals and descriptive captions for the first Audiocaption task, released by RicherMans in 2024. It serves as a benchmark for automated audio description systems and includes baseline code for performance evaluation.
LVBench is a benchmark for long video understanding featuring videos up to two hours in duration, released by zai-org in June 2024. It contains approximately 1,000 records designed to evaluate multimodal models on visual question answering and multiple-choice tasks. The dataset addresses the challenge of extracting information from extended temporal windows that exceed standard video benchmarks.
Multifaceted Collection is a dataset for aligning large language models to diverse human preferences, using system messages to represent individual preferences. The dataset was created by KAIST AI and released in June 2024. Instructions are sourced from five existing datasets.
Hindi VQA is a dataset for visual question answering in Hindi. It was filtered to be more balanced and processed to create sentence embeddings using a pre-trained transformer model, followed by KMeans clustering and t-SNE for visualization. The dataset was uploaded by damerajee to Hugging Face on June 2, 2024.
RadFM_data_csv is a collection of files used for training and testing the RadFM foundation model. The dataset includes a radiology test set with captions and article links, a visual question-answering subset for radiology images, and linked article contents. It was authored by chaoyi-wu and last updated on 2024-06-02.
Four categories of block diagram imagesβBD-EnKo, CBD, FC_A, and FC_Bβare referenced, though only the BD-EnKo subset is provided for summarization research. It facilitates the study of local-global fusion for visual-textual integration as presented at ACL 2024.
RLHF-V-Dataset is a large-scale multimodal feedback dataset constructed using open-source models for reinforcement learning. It was released by the openbmb organization in May 2024 and has been utilized in models like MiniCPM-V 2.0. The dataset is designed for diverse tasks involving computer vision and large language models.
CommonCatalog CC-BY provides approximately 100 million high-resolution images paired with synthetic captions, released by common-canvas in 2024. The collection originates from Yahoo Flickr data from 2014 and features images with resolutions up to 4k.
Over 55,000 real-world user and LLM conversations with associated user preferences, collected from battles between over 70 state-of-the-art LLMs. It was created for a Kaggle competition to predict human preferences in chatbot responses.
A French translation of the Anthropic HH-RLHF dataset, created to support alignment research in the French NLP community. The dataset was uploaded by AIffl and last updated on June 15, -2024. Its specific size, row count, and column structure are not detailed in the provided metadata.