Over 360,000 document images from the PubMed Central Open Access Subset with layout annotations. It provides both bounding boxes and polygonal segmentations for document elements, generated by programmatically matching PDF visual layouts with XML structural data.
Use Cases
- Train object detection models to recognize document structures using the bounding box annotations
- Develop instance segmentation algorithms for scientific papers using the polygonal segmentation masks
- Validate PDF parsing tools by comparing extracted layout elements against the XML-matched ground truth
Strengths
- Includes both bounding boxes and polygonal segmentations for document layout elements
- Sourced from the PubMed Central Open Access Subset, making it available for commercial use
- Annotations are automatically generated by matching PDF format and XML format of scientific articles