Persian Ocr Synth Sentences 100K is a text dataset published on the HuggingFace platform by WeightedAI and last updated on October 5, 2025. The title suggests it likely contains 100,000 synthetically generated sentences in the Persian (Farsi) language, potentially designed for optical character recognition tasks. Specific details regarding the data's structure, source, and generation methodology are not provided in the available metadata.
Use Cases
- Train an OCR model to recognize Persian script (inferred from domain, verify after download)
- Benchmark text recognition accuracy on synthetic Persian sentences (inferred from domain, verify after download)
- Generate synthetic training data for Persian NLP pipelines (inferred from domain, verify after download)
Strengths
- Published on the HuggingFace platform, a major repository for machine learning datasets.
- Last updated on 2025-10-05 19:41:51, indicating recent maintenance.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- WeightedAI
- Freshness
- Last updated 2025-10-05 19:41:51.