Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A multimodal dataset for training optical character recognition models on the Khmer language, created by KiteAether and hosted on Hugging Face. It was last updated on June 11, 2025. The dataset is categorized as containing between 1,000 and 10,000 samples.
License information is unknown, which is critical for determining permissible use. The dataset requires tools capable of handling multimodal (image-text) data and Parquet files.