AncientDoc is a benchmark dataset for Chinese ancient document understanding created by yuchuan123 and hosted on Hugging Face. It contains 2,973 pages from approximately 100 documents across 14 literary types, spanning from the Warring States period to the Qing dynasty. The dataset supports multiple tasks including page-level OCR, vernacular translation, and reasoning-based question answering.
Use Cases
- Train OCR models for complex layouts based on page-level images containing vertical text, variant characters, and annotations
- Develop translation models for classical Chinese based on vernacular translation tasks
- Build question-answering systems for implicit reasoning based on reasoning-based QA tasks
- Evaluate knowledge extraction from historical texts based on knowledge-based QA tasks
- Analyze linguistic styles and rhetoric based on linguistic variant QA tasks
Strengths
- 2,973 pages provide a substantial corpus for training and evaluation
- 14 distinct literary types offer diversity in document genres
- Time span from Warring States to Qing dynasty covers multiple historical periods
- Multiple defined tasks (OCR, translation, QA) create a structured benchmark
Limitations
- Column-level documentation is absent; field semantics must be inferred after download
- Row count is unknown, which may limit suitability assessment
- Data may reflect temporal and literary bias inherent to the selected historical documents
Provenance
- Source
- yuchuan123
- Time Range
- Warring States period to Qing dynasty
- Freshness
- Last updated 2025-08-14 11:37:31; freshness should be verified
- Geography
- China