Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Nearly 40,000 books in the Uzbek language form this text corpus, divided into 'original' and 'lat' branches for OCR and Latin versions. Author murodbek released the dataset to support research on low-resource languages. It was last updated on Hugging Face in June 2025.
Users must refer to the Hugging Face dataset page for loading scripts and full description, as key details like license are not provided here.