Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Infinity-MM provides between 10 million and 100 million multimodal instruction samples, developed by the Beijing Academy of Artificial Intelligence (BAAI) in 2024. It utilizes a synthetic generation pipeline to create detailed image annotations and diverse question-answer pairs for bilingual model training in English and Chinese.
The dataset is released under the CC BY-SA 4.0 license; users should refer to Arxiv paper 2410.18558 for specific details on the synthetic generation methodology.