Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
DocSynth300K provides 300,000 document layout records for large-scale model pre-training, totaling 113GB in size. Released by researcher juliozhao in October 2024, the dataset is designed to enhance model performance in document layout analysis tasks. It is distributed in Parquet format and is compatible with high-performance data libraries like Polars and Dask.
Users must use the huggingface_hub snapshot_download command to retrieve the 113GB dataset; the data is optimized for use with the Polars and Dask libraries.