PubTables-1M_OTSL-v1.1 is a filtered version of the PubTables-1M dataset, containing tables enriched with header information. The dataset is designed for evaluating object detection models and image-to-text methods. It was created by docling-project and last updated on 2025-02-10.
Use Cases
- Evaluating object detection models based on annotated table structures.
- Training image-to-text methods based on the enriched header information.
- Benchmarking table extraction pipelines on a filtered subset of the original PubTables-1M corpus.
Strengths
- Dataset is specifically designed for evaluating both object detection and image-to-text methods.
- Based on the published work 'PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents'.
- Last updated on 2025-02-10, indicating recent maintenance.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Row count is unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Source
- docling-project
- Collection Method
- Filtered version of the original PubTables-1M dataset.
- Freshness
- Last updated 2025-02-10 12:22:27.