Five categories of text areas including "key", "value", "header", "other", and "background" are annotated across document images in this revised version of the FUNSD dataset. The data focuses on correcting connectivity inconsistencies between text areas to better support key-value extraction tasks.
Use Cases
- Train a key-value pair extraction model using the connectivity relations between "key" and "value" labels.
- Develop document classification algorithms based on the distribution of "header" and "other" text areas.
- Evaluate the performance of layout parsing models using the corrected ground truth for document image text areas.
Strengths
- Contains text area annotations for "key", "value", "header", "other", and "background" categories.
- Provides corrected connectivity links representing key-value relations between document segments.
- Addresses specific labeling issues identified in the original FUNSD dataset to improve applicability for information extraction.