Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Over 140,000 English books from the Library of Congress, comprising approximately 8 billion words, are included in this public domain collection. The dataset was compiled by Sebastian Majstorovic via the LoC JSON API, focusing on the Selected Digitized Books collection. It was last updated on the Hugging Face platform in March 2024.
License is listed as unknown on the platform, requiring verification before commercial use. The dataset page notes OCR texts, implying potential noise in the digitized content.