Loading...
Loading...
Image-text pairs, instruction tuning, visual QA, cross-modal data, foundation model training data
1,947 datasets
AYI-NEDJIMI's dataset covers the open source large language model value chain from fine-tuning to production deployment. The description suggests it serves as a technical reference for mastering techniques like LoRA, QLoRA, DPO, RLHF, GPTQ, GGUF, and AWQ. Last updated on February 13, 2026, its specific content and scale require inspection via the linked Hugging Face page.
February 21, 2026 marks the creation of this dataset by Willy08. It contains 11 carefully selected examples of blind spots discovered while experimenting with the Nanbeige/Nanbeige4-3B-Base model. The examples are deliberately diverse and target real weaknesses that even frontier models showed in 2026.
MMSI-Video-Bench is a holistic benchmark for evaluating spatial intelligence in video-based multimodal models. The dataset, created by author 'rbler', includes video clips and was last updated on February 10, 2026. It is hosted on Hugging Face and has been integrated into the VLMEvalKit framework.
A counterfactual VQA dataset constructed using CLEVR blender assets to procedurally generate both negative and normal counterfactual images and questions. The dataset was created by author 'scholo' for the Multimodal Benchmark paper and was last updated on Hugging Face in February 2026. It contains original images, counterfactual variants, and corresponding questions.
SJTU-ViSYS developed M2DGR, a multi-modal and multi-scenario dataset for ground robot navigation, published in RA-L 2021 and ICRA 2022. It provides synchronized sensor data across diverse environments to support Simultaneous Localization and Mapping (SLAM) research.
A collection of model checkpoints for a vision-language model, published on Kaggle. The specific architecture, training data, and performance metrics are not detailed in the available metadata. The author, organization, and last update date are unknown.
Pre-rendered 3D multi-room environments support the Theory of Space benchmark for evaluating spatial reasoning in Vision Language Models. The dataset is designed to test whether foundation models can construct spatial beliefs through active exploration. It was created by MLL-Lab and last updated on February 11, 2026.
A dataset titled 'nexus-hh-rlhf-enriched' published on Kaggle. The title suggests it contains data enriched for Reinforcement Learning from Human Feedback (RLHF), likely involving human preferences for language model outputs. Specific details on size, origin, and creation date are unavailable from the provided metadata.
Kaggle hosts this dataset on power-grid worker safety behavior. The raw description indicates it contains multimodal data related to risk and standard operating procedure (SOP) operations. The dataset's author, organization, and specific scale are unknown.
A dataset likely containing multiple data types related to phishing attacks. The dataset is published on Kaggle, but its specific contents, size, and creation details are not described. Further verification after download is required to confirm its scope and utility.
RadImgNet-VQA is a dataset hosted on Kaggle, likely designed for visual question answering tasks in the medical domain. The title suggests it contains pairs of radiology images and associated questions, potentially for training AI models to interpret medical scans. Its specific size, source, and creation date are not provided in the available metadata.
NVIDIA's PhysicalAI dataset provides pre-processed 3D assets for predicting volumetric mechanical properties. The dataset combines four individual 3D asset collections, processed to include multi-view renders, voxelized representations, and LLM-annotated material descriptions. It was last updated on February 5, 2026.
Tokyo driving data provides a large-scale visual question answering dataset for physically grounded spatiotemporal reasoning. It contains 16 million question-answer pairs over 270,000 frames, constructed from 100 hours of multi-sensor driving data. The dataset was created by turing-motors and last updated on the platform in January 2026.
Multimodal_Diet_Dataset is a dataset hosted on Kaggle. Its title suggests it contains data related to diet and nutrition, potentially combining multiple data types. Further details regarding its size, origin, and specific contents are unavailable from the provided metadata.
WorldVQA is a benchmark dataset created by MoonshotAI to evaluate atomic vision-centric world knowledge in Multimodal Large Language Models (MLLMs). It was last updated in February 2026. The dataset decouples visual knowledge retrieval from reasoning to provide a strict measurement of a model's fundamental world knowledge.
Blip3O 256 is a dataset authored by diffusion-bench and hosted on Hugging Face. The dataset was last updated on March 25, 2026. Its specific content and scale are not detailed in the available metadata.
A longitudinal and multimodal benchmark for robust drift detection in Android malware. The dataset is hosted on Kaggle, but specific details on its size, creation date, and authorship are not provided in the available metadata. Its primary purpose is to serve as a testbed for evaluating the robustness of machine learning models against concept drift in the malware domain.
Testing-multimodal is a dataset published on Kaggle. The title suggests it is intended for evaluating machine learning models that process multiple data types. The dataset's specific content, size, and origin are not detailed in the available metadata.
SpaVis-6M is a multimodal dataset for computational pathology, integrating visual and molecular data. It was created by minghaofdu and is associated with research presented at ICLR 2026. The dataset page was last updated on February 12, 2026.
122,000 vision-question-answer pairs across more than 145 microscopy genera. The dataset likely contains images paired with textual questions and answers for visual question answering tasks. Published on Kaggle.