Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,956 datasets
Open-ended questions and images are the primary categories in this multimodal dataset. These samples require the integration of vision, language, and commonsense knowledge for successful completion.
Image Captions is a multimodal dataset hosted on HuggingFace by csarron, last updated in November 2021. It pairs images with descriptive text, facilitating tasks that link visual and language understanding. The specific number of image-text pairs and source of the images are not detailed in the available metadata.
10,921 high-resolution remote sensing images collected from satellite imagery sources, each paired with 5 descriptive natural language captions. The dataset covers 30 distinct scene categories, including airports, bridges, and residential areas, totaling approximately 54,605 caption-image pairs.
This repository provides a PyTorch implementation for deep learning cross-modal hashing. It was authored by WangGodder and last updated in October 2021. The specific dataset details, including row and column counts, are unknown.
The Wikipedia-based Image Text (WIT) Dataset contains 37.6 million entity-rich image-text examples paired with 11.5 million unique images across 108 Wikipedia languages. It was created by keshan for pretraining multimodal machine learning models and was last updated in August 2021.
Encompassing biomechanical data from a study investigating how flamingos support their body on one leg with minimal muscle force. The research includes measurements from both cadaveric specimens and live flamingos, analyzing body sway and joint posture. The dataset is associated with a study published in 2020 by author Young-Hui Chang.
103 participants, including 30 homosexual men, 35 heterosexual men, and 38 heterosexual women, underwent multimodal MRI scans. Amirhossein Manzouri published this study in 2020, comparing cortical thickness, surface area, subcortical volumes, and resting-state functional connectivity across groups.
ENVIDAT presents stated preference data on improved forest management measures from seven Swiss municipalities in Grisons and Valais. The data was collected via an online questionnaire between October 2019 and February 2020, receiving 939 responses from 10289 invited households. It includes a choice experiment with twelve tasks assessing willingness to pay for avalanche and rock fall risk reduction, alongside sociodemographic and attitudinal questions.