Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Holding approximately 200,000 image-text pairs for LaTeX optical character recognition, featuring both printed and synthetic handwritten mathematical formulas. Developed by linxy and updated in December 2024, the dataset aggregates data from Zenodo, CROHME, and custom-built sources to support the LaTeX_OCR project.
The dataset is distributed in Parquet format under the Apache 2.0 license. Users should note that the synthetic handwritten set is derived directly from the formulas in the printed set.