Loading...
Loading...
3,021 annotated document page images form the training, evaluation, and test sets for the ICDAR 2019 Competition on Baseline Detection (cBAD). The dataset was created by Markus Diem of TU Wien and consists of real-world images collected from seven European archives, with all baselines manually annotated. The training and evaluation sets contain PAGE XML files with annotated text regions and baselines.
The ground truth for the test set was published after the competition deadline (May 2019).