Approximately 1.8 million synthetically generated CAPTCHA images created by szili2011. Each image contains a random, case-sensitive sequence of letters (a-z, A-Z) and numbers (0-9). The dataset was last updated on June 26, 2025.
Use Cases
- Train OCR models to recognize alphanumeric characters based on the synthetic CAPTCHA images.
- Benchmark the performance of character recognition algorithms on case-sensitive sequences.
- Generate adversarial examples to test the robustness of vision models against distorted text.
- Develop data augmentation pipelines for text-in-the-wild recognition tasks.
Strengths
- Contains approximately 1.8 million images, providing a large-scale resource for training.
- Images are synthetically generated, allowing for controlled variation and reproducibility.
- Specifically built for OCR training and testing, as stated in the description.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Data may reflect bias inherent to the synthetic generation method used.
Provenance
- Source
- szili2011 on Hugging Face
- Collection Method
- Synthetically, artificially generated.
- Time Range
- null
- Freshness
- Last updated 2025-06-26 15:27:29; freshness should be verified.
- Geography
- null