Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,947 datasets
KOL Decision-Making Dataset for Web3 Community Managers is designed for instruction tuning. The dataset likely contains examples of decisions or actions taken by Key Opinion Leaders in Web3 communities. Its origin and scale are unspecified, as the description metadata is limited.
Double-Delta Multi-Fidelity Aerodynamics Dataset is an open-source benchmark for a parametric family of double-delta wings. It contains paired low-fidelity Vortex Lattice Method and high-fidelity simulation data. The dataset was created by yirens and last updated on 2026-03-22.
A dataset for image captioning tasks, likely containing images paired with descriptive text in the Kannada language. It is hosted on the Kaggle platform, but details about its size, creation date, and authorship are not provided in the available metadata. The dataset's content and structure require verification after download.
Kaggle hosts a dataset titled 'streamvlm-checkpoint'. The dataset likely contains model weights or parameters for a vision-language model. No information is available regarding its author, organization, size, or last update date.
NOO-Verified-Global-Entities provides a data infrastructure layer to prevent AI agents from hallucinating non-existent or unsuitable B2B suppliers. The dataset was created by Nooxus-AI and was last updated in March 2026. It is designed as a definitive verification source for global commercial entities.
DatapointAI released a 1,000-row dataset in March 2026 for evaluating image-to-video generation models. Each row contains a reference image, two generated videos from Pika and CogVideoX models, and 10 aggregated human preference annotations. The dataset provides a total of 10,000 individual human judgments on video quality.
A dataset focused on dance movement and pose analysis, published on Kaggle. The dataset likely contains motion capture or video data paired with pose annotations. Specific details on size, collection method, and temporal coverage are not provided in the available metadata.
3,000 rows of human preference data for evaluating image-to-video generation. Each row contains a reference image, two generated videos from Pika and CogVideoX models, and 10 aggregated human annotations. The dataset was created by datapointai and last updated in March 2026.
ZwZ-RL-VQA is a dataset containing 74,000 high-quality visual question-answering pairs generated via Region-to-Image Distillation. The dataset was created by inclusionAI for training multimodal large language models on fine-grained perception tasks and was last updated in March 2026.
BovCap-5K provides a collection of cattle images paired with natural language descriptions for research. The dataset's author, organization, and specific scale are not detailed in the provided metadata. Its last update date and licensing terms are also unknown.
A dataset for Visual Question Answering (VQA) tasks in the medical domain. It is hosted on Kaggle, but its specific size, creation date, and authorship are not detailed in the provided metadata. The dataset likely contains pairs of medical images and related questions with answers.
Open Access journal articles up to February 2026 used for domain-adaptive pretraining and instruction tuning of the AdditiveLLM2 model. The dataset includes text and images, and is split by source journal. It was created by ppak10 and last updated on March 25, 2026.
Tables 1-6 from USGS Open-File Report 02-59 contain data on salinity, discharge, and stage (water level) related to culverts under the main road in Everglades National Park. The data were gathered as part of a 2002 study by the South Florida Natural Resources Center and USGS to assess the road's influence on salinity intrusion into Florida Bay. Monitoring sites recorded water level, salinity, and flow during periods when water was present.
Avencast's dataset, associated with the arXiv preprint 'EveNet: A Foundation Model for Particle Collision Data Analysis', was last updated on March 31, III. The dataset appears to be designed for training and evaluating foundation models in the domain of particle physics, specifically for analyzing collision event data.
AEC-Bench is a multimodal collection of real-world Architecture, Engineering, and Construction documents, including construction drawings and floor plans. The dataset was created by nomic-ai and was last updated in April 2026. It is structured for benchmarking tasks across scopes and task families.
Medical VQA - 5 Datasets (LLaVA format) is a collection of medical vision-language datasets aggregated on Kaggle. The datasets are formatted for the LLaVA (Large Language-and-Vision Assistant) framework, suggesting they contain paired image and text data. The specific source, size, and creation date of the datasets are not provided in the available metadata.
A dataset of 2,000 human preference annotations for evaluating image-to-video generation. Each row contains a reference image, two generated videos from Pika and CogVideoX models, and 10 human annotations aggregated via majority vote. Created by datapointai and last updated in March 2026.
MIBench is a benchmark designed to evaluate the multimodal interaction capabilities of Large Multimodal Models (LMMs). It was created by an author or organization named Resurrect and was last updated on March 26, 2026. The benchmark focuses on how models integrate and utilize information across different modalities based on task demands.
Wanglab's training dataset for reinforcement learning optimization of the BioReason-Pro model. The data contains proteins with Gene Ontology term annotations, InterPro domains, STRING protein-protein interactions, and protein metadata. It was last updated on March 20, 2026.
A multimodal dataset for fine-tuning Large Language and Vision Assistant (LLaVA) models. It likely contains medical images and associated text from the HAM10000 and BCN20000 collections. The dataset is published on Kaggle, but specific details like size, format, and update date are unknown.