Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
30 patients with basal cell carcinomas contributed to this multimodal dataset of paired reflectance confocal microscopy images and Raman spectra. The data was collected via point-by-point scanning and is authored by Khan, Fadeel Sher, hosted by the Texas Data Repository. The dataset was last updated on March 18, 2024.
484 webpages from the C4 validation set serve as a testbed for multimodal large language models. The dataset focuses on the task of converting visual designs into code implementations. SALT-NLP created this resource, which was last updated on March 11, 2024.
Gold-standard benchmark for document alignment between Sinhala, Tamil, and English languages. It contains manually annotated document pairs crawled from four Sri Lankan news websites: Army, Hiru, ITN, and Newsfirst.
Presenting a gold-standard benchmark dataset for sentence alignment between Sinhala, English, and Tamil languages. The data was crawled from news websites including Army, Hiru, ITN, and Newsfirst, with aligned sentences derived from a prior document alignment dataset.
A combined dataset for medical visual question answering, merging the VQARAD and SLAKE collections. The dataset was created by Shashwath01 and was last updated on March 5, 2024. It has been used to train a specific model hosted on Hugging Face.
LLM-jp, a collaborative project in Japan, provides this dataset. It is a Japanese translation of a 21,000-instruction English subset from the OASST1 dataset, created using the DeepL translation service. The dataset was last updated on February我们发现一个错误。根据输入,数据集标题是“Oasst1 21K Ja”,描述中提到它是“oasst1-21k-ja”,并说明是“Japanese translation of an English subset of oasst1”。因此,正确的摘要应基于此信息。输入中没有明确的行数“21,000”,但标题和名称暗示了“21k”。我将据此修正摘要。
Human Preference Dataset v2 (HPD v2) is a large-scale collection of human preference choices on images generated by text-to-image models. It contains 798,000 preference choices across 430,000 images. The dataset was created by ymhao and was last updated on February 21,我们发现了一个错误。
AGIEval is a human-centric benchmark for evaluating foundation models. This dataset contains the Chinese-language LogiQA subtask, processed from the Microsoft AGIEval GitHub repository. It was authored by 'hails' and last updated on January 26, 2024.
AGIEval is a human-centric benchmark for evaluating foundation models. This dataset contains the JEC-QA-CA subtask, which likely contains Chinese question-answering data. The dataset was processed from the AGIEval repository by the user 'hails' and was last updated on the Hugging Face platform on 2024-01-26.
ColorSwap is a multimodal dataset of 2,000 unique image-caption pairs, grouped into 1,000 examples. It was created by stanfordnlp and last updated in February 2024. The dataset is designed to assess and improve the proficiency of multimodal models in matching objects with their colors.
A collection of multi-turn guessing games utilizing VisualGenome images and scene graphs for attribute grounding tasks. The data serves as a multi-task framework to evaluate the quality of neural representations through object identification and visual dialogue.
Encompassing 0.5 million synthetic Chinese document images generated by the SynthDoG tool for training the Donut model. It is part of a multi-language collection created by naver-clova-ix and was last updated in January 2024.
VAST is an omni-modality dataset and foundation model from NeurIPS 2023 containing four distinct data categories: vision, audio, subtitles, and text. It provides a framework for multi-modal learning where visual frames are paired with corresponding sound, textual transcripts, and descriptive text.
AGIEval is a human-centric benchmark for evaluating foundation models. This dataset contains the Gaokao Biology subtask, processed from the AGIEval repository. The data was authored by 'hails' and last updated on January 26, 2024.
150,000 GPT-generated multimodal instruction-following data points collected in April 2023. The dataset utilizes the GPT-4-0314 API to synthesize vision-language interactions for the development of large multimodal models.
Between 10,000 and 100,000 expert-annotated sentences comprise this dataset for token-level acronym identification in the scientific domain. Created by Amirveyseh for the AAAI-21 Workshop on Scientific Document Understanding, it includes standardized training, validation, and test splits.
4.8 million collective human preferences compare the helpfulness of two responses to questions or instructions. The dataset spans 129 diverse subject areas, from cooking to legal advice, and is an extended version of the original 385K SHP dataset. Created by stanfordnlp and updated in January 2024, it is intended for training RLHF reward models and NLG evaluation models.
MMVP-VLM is a benchmark dataset designed to systematically evaluate the performance of recent CLIP-based visual language models in understanding and processing visual patterns. It distills a subset of questions from the original MMVP benchmark into simpler language descriptions, categorizing them into distinct visual patterns. The dataset was created by the author 'MMVP' and was last updated on Hugging Face on January 10, 2024.
A-OKVQA is a dataset for visual question answering that requires external knowledge and reasoning. The dataset was created by HuggingFaceM4 and was last updated in February 2024.
12 million image-text pairs sourced from 350 manually curated subreddits covering diverse objects and scenes. The dataset utilizes subreddit names as coarse labels to guide composition without requiring manual per-instance annotation.