OCR text and image descriptions extracted from Thai Astrology PDF documents using the Gemma 4 31B multimodal model. The dataset includes columns for source PDF name, page number, and original page image. The dataset was created by Phonsiri and last updated on May 25, 2026.
Use Cases
- Training OCR models based on Thai PDF text extraction.
- Developing multimodal AI systems based on paired images and captions.
- Analyzing Thai astrology content based on extracted text and descriptions.
- Benchmarking large multimodal models on document understanding tasks.
Strengths
- Extraction performed using the Gemma 4 31B multimodal model.
- Dataset structure includes source PDF name, page number, and original page image.
Limitations
- Row count, file formats, and license information are unknown.
- Column-level documentation beyond the three listed fields is absent; field semantics must be inferred after download.
- Freshness should be verified as the last update date is May 25, 2026.
Provenance
- Source
- Thai Astrology PDF documents.
- Collection Method
- OCR text and image descriptions extracted using the Gemma 4 31B multimodal model.
- Freshness
- Last updated 2026-05-25 10:47:13
- Geography
- Thailand (likely, based on source content)