Sign in to view source links and access this dataset
Description
Francisco-Cruz provides 1,003 images of invoices and receipts, each with transcriptions for key financial fields. The dataset includes annotations for seller details, tax IDs, dates, and amounts. It was last updated on Hugging Face in May 2024.
Use Cases
Train optical character recognition (OCR) models based on the provided invoice and receipt images.
Develop models for structured information extraction based on annotated fields like seller name, address, and tax IDs.
Benchmark document understanding systems for financial documents using the provided JSON annotations.
Fine-tune models for invoice date and total amount recognition from scanned documents.
Strengths
Contains 1,003 images, providing a substantial corpus for model training.
Includes structured JSON annotations for eight specific financial fields per document.
Limitations
Row count and file formats for the underlying data are unknown, which may limit suitability assessment.
Column-level documentation is absent; field semantics must be inferred after download.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
huggingface
Collection Method
Collection method is not specified in the provided description.
Time Range
Temporal coverage is not specified in the provided description.
Freshness
Last updated 2024-05-02 22:50:39; freshness should be verified.
Geography
The dataset title and author name suggest a potential Portuguese (PT) geographic focus.
License is unknown; users must verify licensing terms before use.