tw-OCR is a dataset for optical character recognition tasks focused on Traditional Chinese. It contains images of real documents from Taiwan, such as forms, announcements, official documents, and academic materials, with annotations. The dataset was curated by Huang Liang Hsun and shared by Twinkle AI under a CC BY-SA 4.0 license.
Use Cases
- Train OCR models based on diverse Traditional Chinese document images.
- Evaluate OCR model performance on real-world Taiwanese documents.
- Benchmark OCR systems on varied fonts, resolutions, and noise conditions mentioned in the description.
- Develop document digitization tools for Taiwanese administrative or academic contexts.
Strengths
- Focuses on Traditional Chinese (zh-tw), a specific linguistic context.
- Includes real document images from common Taiwanese scenarios.
- Covers diverse fonts, resolutions, and noise conditions as stated in the description.
Limitations
- Row count, file formats, and column-level documentation are unknown.
- Dataset size is unspecified, limiting suitability assessment.
- Freshness should be verified as the last metadata update was 2025-05-21.
Provenance
- Source
- Twinkle AI
- Collection Method
- Curated from real documents common in Taiwan.
- Freshness
- Last updated 2025-05-21 02:16:57.
- Geography
- Taiwan