1,930 train ticket images categorized into 1,530 synthetic and 400 real-world samples for document information extraction. The dataset captures eight specific text fields such as ticket number, train number, and seat category across Chinese, English, and numeric scripts.
Use Cases
- Train a field extraction model to identify 'starting station' and 'destination station' within a fixed layout
- Evaluate OCR robustness against imaging distortions using the 80 real-world test images
- Develop a text recognition system for mixed-script 'ticket rates' and 'seat category' fields
Strengths
- 1,530 synthetic and 400 real-world images (320 train / 80 test)
- Eight labeled text fields: ticket number, starting station, train number, destination station, date, ticket rates, seat category, and name
- Multilingual support for digits, English characters, and Chinese characters