Sign in to view source links and access this dataset
Description
MMAT-1M is a million-scale multimodal agent tuning dataset built by consolidating subsets of five publicly available datasets: Visual CoT, LLaVA-CoT, The Cauldron, TabMWP, and Infoseek. It integrates dynamically generated API calls and Retrieval Augmented Generation (RAG) information through a GPT-4o-powered multi-turn paradigm. The dataset was created by VIS-MPU-Agent and was last updated on August 4, 2025.
Use Cases
Fine-tuning multimodal AI agents based on the integrated API calls and RAG information.
Training models for complex question-answering based on the multi-turn paradigm described.
Benchmarking agent performance on tasks requiring rationales refined via the described process.
Developing instruction-following models using the consolidated question-answer datasets.
Strengths
Million-scale size, as indicated by the dataset name.
Integrates five distinct source datasets, suggesting breadth.
Includes dynamically generated API calls and RAG information.
Uses a GPT-4o-powered multi-turn paradigm for data construction.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is specified only as 'million-scale'; exact size is unknown.
Data may reflect biases inherent to the five source datasets from huggingface.
Provenance
Source
VIS-MPU-Agent via Hugging Face.
Collection Method
Consolidated from subsets of Visual CoT, LLaVA-CoT, The Cauldron, TabMWP, and Infoseek datasets, with integration of API calls and RAG via GPT-4o.
Freshness
Last updated 2025-08-04 02:50:54.
License is unknown; users must verify permissions before use.