A conversion of the original SynthTabNet dataset into the OTSL format, as presented in the paper 'Optimized Table Tokenization for Table Structure Recognition'. The dataset includes original annotations and new additions, organized into 4 parts totaling 600,000 tables with varied appearances. It was created by docling-project and last updated on Hugging Face in August 2023.
Use Cases
- Training table structure recognition models based on synthetic tables with varied appearances
- Benchmarking tokenization methods for tables based on the OTSL format
- Developing document processing pipelines based on a large-scale synthetic table corpus
- Studying the impact of table size, structure, style, and content on recognition performance
Strengths
- Contains 600,000 tables, providing a large-scale synthetic corpus
- Organized into 4 parts of 150,000 tables each, with varied appearances
- Includes both original annotations and new additions
Limitations
- Column-level documentation is absent; field semantics must be inferred after download
- Row count is unknown, which may limit suitability assessment
- Last updated 2023-08-31 17:14:02; freshness should be verified
Provenance
- Source
- docling-project
- Collection Method
- A conversion of the original SynthTabNet dataset into the OTSL format.
- Freshness
- Last updated 2023-08-31 17:14:02