14 million scene text images organized into 11 subsets representing diverse real-world challenges like curved and artistic text. The collection serves as a massive-scale pre-training corpus for Scene Text Recognition (STR) models to improve performance on irregular text.
Use Cases
- Pre-train vision-language models for OCR using the 14 million image-text pairs
- Benchmark model performance on irregular text geometries using the 'Curved' and 'Multi-oriented' subsets
- Improve text recognition in degraded conditions by training on the 'Blurry' and 'Low-resolution' data categories
- Develop text detection and recognition pipelines using the provided bounding box and transcription labels
Strengths
- 14 million images across 11 distinct subsets for large-scale STR pre-training
- Includes 400,000 manually annotated images for a challenging evaluation benchmark
- Covers specific text categories including 'Curved', 'Multi-oriented', 'Artistic', and 'Low-resolution'
- Provides image-level text transcriptions and bounding box coordinates for all samples