VietHTR-Line is a dataset of 60,247 Vietnamese handwritten text images curated by 5CD-AI for Handwritten Text Recognition research. The images were crawled from public internet sources. The dataset was last updated on July 22, -2026.
Use Cases
- Training Vietnamese OCR models based on the collection of handwritten text images.
- Benchmarking the performance of handwriting recognition systems on a specific language dataset.
- Developing data augmentation techniques for handwritten text based on a large image corpus.
- Studying the characteristics of Vietnamese handwriting styles based on images from public sources.
Strengths
- Contains 60,247 individual handwritten text images, providing a substantial corpus for model training.
- Specifically focused on the Vietnamese language, addressing a need for local, high-quality data.
- Explicitly curated for Handwritten Text Recognition research, indicating a clear purpose.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- The description metadata is limited; actual data quality, such as annotation consistency, requires manual inspection.
- The original source is described as 'public internet sources', which may introduce unknown biases in handwriting styles or content.
Provenance
- Source
- 5CD-AI
- Collection Method
- Images were crawled from public internet sources.
- Time Range
- null
- Freshness
- Last updated 2026-07-22 23:46:39; freshness should be verified.
- Geography
- Vietnam