Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A collection of 9.3 million document images processed through Optical Character Recognition (OCR) using the docTR library. The dataset is derived from the AMF-PDF dataset, part of the Finance Commons collection, and was created by lightonai. The dataset card was last updated on September 23, 2024.
The dataset page notes that native text annotations in a related dataset have imperfections; users should verify OCR quality. The full description is hosted externally.