Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Between 1 million and 10 million Japanese-translated vision-language records comprise this collection created by turing-motors in 2024. It adapts the 50-dataset Cauldron collection used for Idefics2 fine-tuning into Japanese using the DeepL API, specifically targeting visual question answering tasks.
This dataset is a machine-translated version of the original English Cauldron dataset; users should verify translation accuracy for sensitive applications. It is provided in Parquet format.