Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Featuring 0.5 million synthetic Korean document images generated by the SynthDoG tool for training the Donut (OCR-Free Document Understanding Transformer) model. It was created by naver-clova-ix and last updated in January 2024.
Users should review the associated GitHub repository (https://github.com/clovaai/donut) and the original paper for details on the SynthDoG generation process and the Donut model. License information is not provided in the input.