A large-scale multilingual document OCR dataset containing approximately 400GB of images with annotations across multiple global languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing. It was created by Nayana-cognitivelab and last updated on 2025-07-21.
Use Cases
- Train multilingual OCR models based on the annotated document images.
- Benchmark document layout analysis performance across languages.
- Develop document understanding systems for languages like Arabic, German, Russian, Spanish, and French.
Strengths
- Approximately 400GB of annotated document images provides a substantial volume for training.
- Multilingual annotations across languages like Arabic, German, Russian, Spanish, and French support cross-lingual work.
- WebDataset format using TAR archives is designed for efficient streaming and processing.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- Nayana-cognitivelab
- Freshness
- Last updated 2025-07-21 12:44:25; freshness should be verified.
- Geography
- Global