Loading...
Loading...
A dataset of 3,021 annotated document page images collected from seven European archives for the ICDAR 2019 Competition on Baseline Detection (cBAD). All baselines were manually annotated, and the training and evaluation sets include PAGE XMLs with annotated text regions and baselines. The dataset was created by Markus Diem of TU Wien.
License is listed as Open Access (green); specific terms should be verified.