Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
OpenGVLab's Doc-750K dataset, referenced in the paper 'Docopilot: Improving Multimodal Models for Document-Level Understanding', is a collection of documents for training AI models. The dataset was last updated on July 22, 2025. It appears to contain a large number of document images, as suggested by unzipping instructions for image archives.
The description notes potential issues when unzipping the image archive on Linux, such as zip bomb warnings or bad zipfile offset errors.