Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
ScriptNet's cBAD dataset contains the training and test set for the ICDAR 2017 Competition on Baseline Detection in Archival Documents. It comprises 2035 annotated document page images collected from 9 different archives, split into two tracks: one for simple handwritten paragraphs and another for complex documents with tables, marginalia, and noise. The training data includes PAGE XML files with manually annotated text regions and baselines.
License is listed as Open Access (green). The test set consists of images only, without annotations.