SEC-EDGAR is a dataset of all major filings from the SEC EDGAR database, containing 590 GB of data spanning 8 million samples and 43 billion tokens. It was released by Datamule, Teraflop AI, and Eventual, with the bulk data collected using the datamule-python library and official API created by John Friedman. The dataset was last updated on 2026-04-08.
Use Cases
- Train large language models based on the 43 billion tokens of financial and legal text.
- Analyze corporate disclosure trends based on the collection of all major SEC filings.
- Develop financial document classification or information extraction systems based on the structured filings.
- Benchmark NLP model performance on domain-specific legal and financial language.
Strengths
- Contains 590 GB of data, providing substantial volume for model training.
- Spans 8 million samples, offering a large number of individual data points.
- Includes 43 billion tokens, a significant scale for text-based AI applications.
- Covers all major filings from the authoritative SEC EDGAR database.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Freshness should be verified as the last update date is in the future (2026-04-08).
Provenance
- Source
- SEC EDGAR database
- Collection Method
- Collected using the datamule-python library and official datamule API created by John Friedman.
- Time Range
- null
- Freshness
- Last updated 2026-04-08 02:29:00
- Geography
- null