A corpus of line-level text images paired with word-segmented Burmese text labels for optical character recognition. The dataset format pairs image file paths with text labels using a tab delimiter. The dataset was created by LULab and last updated on December 20, 2024.
Use Cases
- Train optical character recognition models based on paired image-text data.
- Benchmark OCR performance for the Burmese language.
- Develop post-OCR error correction techniques based on the described dataset.
- Study word segmentation in Burmese text using the underscore delimiter mentioned.
Strengths
- Text labels are specifically word-segmented with an underscore delimiter.
- Data is structured as paired image file paths and text labels.
- The dataset is associated with a published research paper cited in the description.
Limitations
- Row count and dataset size are unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- LULab
- Freshness
- Last updated 2024-12-20 10:57:15; freshness should be verified.