500 annotated invoice documents formatted specifically for fine-tuning the Donut document understanding model. The data originates from the Mendeley 'Samples of electronic invoices' collection and was processed by the Katana ML team for the Sparrow open-source project.
Use Cases
- Fine-tune a Donut model for end-to-end invoice parsing using the provided ground truth annotations.
- Evaluate document image-to-text transformer performance on the 500 electronic invoice samples.
- Test data extraction pipelines within the Sparrow open-source framework using the processed invoice images.
Strengths
- 500 annotated invoice documents processed for machine learning tasks.
- Formatted specifically for the Donut (Document Understanding Transformer) model architecture.
- Derived from the peer-reviewed Kozłowski and Weichbroth (2021) Mendeley Data repository.
- Includes ground truth annotations required for training document-to-text models.