Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
A dataset likely containing human preference data for Reinforcement Learning from Human Feedback (RLHF) applications. It was published by author liyucheng on the Hugging Face platform on April 15, 2023. The dataset's title suggests a connection to the Chinese Q&A platform Zhihu and a scale of approximately 3,000 entries.
A dataset of annotations for Visual Question Answering (VQA) tasks based on the COCO dataset. The dataset was uploaded by author 'ryanramos' to the Hugging Face platform and was last updated on April 3, 2023. The specific number of rows, column structure, and license information are not provided in the available metadata.
OK-VQA_train is a dataset for visual question answering tasks, likely containing image-question-answer pairs. The dataset was created by Multimodal-Fatima and was last updated on March 23, 2023. Specific details on the number of samples, data format, and license are not provided in the available metadata.
An image-text dataset likely containing pictures of nail sets. The dataset was published by Boyuan07 on HuggingFace and was last updated on March 29, 2023. The specific content, scale, and structure require verification after download.
5,000 test images from the MSCOCO 2014 collection paired with human-annotated captions for image-text retrieval tasks. The data follows the Karpathy split, a standard benchmark for evaluating cross-modal alignment between visual features and natural language descriptions.
A blind evaluation dataset of high-quality, diverse, human-written instructions with demonstrations. The dataset was created by HuggingFaceH4 and last updated on February 28, 2023. It is intended for use in step 3 evaluations within a Reinforcement Learning from Human Feedback pipeline.
Anthropic's HH dataset reformatted into prompt, chosen, and rejected samples by Dahoas. The data was last updated on Hugging Face in February 2023. It provides a structured format for training and evaluating language models using human preferences.
32 features across 5 categories like Environment and Damage annotate public videos of natural disasters. The dataset was used for the TRECVID DSDI task from 2020-2022 and is maintained by the National Institute of Standards and Technology. All footage consists of airborne, low-altitude video from disaster events.
Cuad VQA is a dataset for visual question answering tasks, published on the Hugging Face platform by user munish0838. The dataset was last updated on March 1, 2023. Its specific content, size, and structure are not detailed in the available metadata.
A subset of 2.66 million Chinese image-text pairs extracted from the LAION-5B-high-resolution multilingual multimodal dataset. The dataset was created by 'wanng' and last updated on Hugging Face in December 2022. The provided metadata file is approximately 381 MB and contains text information such as URLs, but does not include the actual image files.
100 million Chinese image-text pairs form a subset of the Noah-Wukong multimodal dataset. The dataset was uploaded by author 'wanng' to Hugging Face and last updated on December 11, 2022. The text metadata for these pairs occupies approximately 16GB of space.
Rlhf Reward Datasets are collections for training reward models in reinforcement learning from human feedback. The dataset was authored by yitingxie and last updated on Hugging Face in January 2023. Specific details on size, format, and content are not provided in the available metadata.
Facebook AI created this benchmark dataset to measure progress in multimodal reasoning for hate speech detection. The dataset pairs images with text to form memes, each labeled for hateful content. It was published in 2020 and last updated on the platform in December 2022.
Sft Hh Rlhf is a dataset published on Hugging Face by Dahoas, with its last update recorded on 2022-12-22. The title suggests it contains data related to reinforcement learning from human feedback (RLHF) and supervised fine-tuning (SFT). The dataset's specific content, scale, and structure require verification after download.
Comprising over 37 million image-text associations extracted from Wikipedia articles. It is a multilingual dataset covering 108 languages, curated by Google Research and released by Wikimedia.
An enriched version of the SROIE 2019 dataset adds labels for line descriptions and line totals to aid OCR and layout understanding. The training split contains 652 samples, each pairing an image with OCR text data. Arvindrajan92 published this multimodal dataset on HuggingFace in October 2022.
Dahoas published a dataset of synthetic prompts on Hugging Face in December 2022. The dataset likely contains text prompts generated for Reinforcement Learning from Human Feedback (RLHF) workflows. Its specific size, format, and column structure are not detailed in the available metadata.
svjack created this dataset to train a PokΓ©mon text-to-image model. It pairs PokΓ©mon images from the FastGAN project with captions generated by the BLIP model, adding a Chinese translation column. The dataset was last updated on Hugging Face in October 2022.
Aggregating an export of all XKCD comics, including their transcript and explanation scraped from explainxkcd.com. It includes fields such as comic title, image URL, transcript, and explanation URL.
Multilingual image captions with annotations and language labels created by experts. It supports numerous languages, including Languageadq, Languageaeu, and Languageabc. The specific number of rows, columns, and file formats is not provided.