A dataset for post-OCR error correction in the Sanskrit language, generated using the RoundTripOCR technique. It was created by cfilt and uploaded to HuggingFace on December 8, 2024. The dataset includes train, test, and validation splits.
Use Cases
- Train OCR post-correction models based on the RoundTripOCR technique.
- Benchmark error correction algorithms for Sanskrit text.
- Evaluate the performance of OCR systems on historical or classical scripts.
- Develop language-specific NLP tools for Sanskrit based on corrected text data.
Strengths
- Dataset includes structured train, test, and validation splits.
- Specifically targets the Sanskrit language, a niche domain.
- Methodology is documented via a public GitHub repository.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- cfilt
- Collection Method
- Generated using the RoundTripOCR technique.
- Freshness
- Last updated 2024-12-08 14:42:34; freshness should be verified.