Over 439,000 unique images and 581,000 image-text pairs were automatically extracted from more than 20,000 PDF documents in the REGIS collection. The dataset, created by Geologi, focuses on technical documents, theses, and reports from the Oil & Gas and Geosciences domain. It was last updated on 2026-06-21.
Use Cases
- Train multimodal vision-language models based on domain-specific image-text pairs.
- Fine-tune image captioning models based on technical figures and their associated captions.
- Develop cross-modal retrieval systems for technical documents based on the paired image and text data.
- Analyze visual patterns in Oil & Gas and Geoscience reports based on the large collection of extracted images.
Strengths
- Large scale with over 439,000 unique images.
- Substantial number of image-text pairs (581,000).
- Sourced from a focused domain collection of over 20,000 PDFs.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- REGIS collection (technical documents, theses, and reports from Oil & Gas and Geosciences).
- Collection Method
- Automatically extracted from PDF documents.
- Freshness
- Last updated 2026-06-21 19:00:39; freshness should be verified.