The dataset contains 8,500 rows, each representing the full text of a single Arabic book, extracted using the arabic-large-nougat model. It spans approximately 1.1 billion tokens, as calculated by the GPT-4 tokenizer. The dataset was authored by MohamedRashad and last updated on 2024-11-28.
Use Cases
- Training large language models for Arabic based on the 1.1 billion tokens of book text.
- Benchmarking Arabic OCR model performance based on texts extracted via the arabic-large-nougat model.
- Conducting linguistic analysis of Arabic literature based on the full-text content of books.
- Fine-tuning text generation models for Arabic based on the book corpus.
Strengths
- Contains 8,500 full-text Arabic books, providing substantial content for analysis.
- Demonstrates a specific extraction method using the arabic-large-nougat model for Arabic OCR.
- Represents a large-scale text corpus of approximately 1.1 billion tokens.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Source
- huggingface
- Collection Method
- Text extracted using the arabic-large-nougat model for Arabic OCR.
- Time Range
- null
- Freshness
- Last updated 2024-11-28 17:47:50; freshness should be verified.
- Geography
- null