Approximately 380 billion tokens of government text and data aggregated from open data programs. The Open Government dataset is curated by AgentPublic and was last updated on January 31, 2025. It currently features collections from the US, France, European, and international organizations, structured through the Finance Commons and Legal Commons initiatives.
Use Cases
- Train large language models on public sector text based on the described 380B token corpus.
- Analyze legal or regulatory language patterns based on the 'Legal Commons' collection mentioned in the description.
- Study financial disclosures and reports based on the 'Finance Commons' collection mentioned in the description.
- Conduct comparative policy analysis across US, French, and international government sources as indicated by the geographic focus.
Strengths
- Large scale of approximately 380 billion tokens.
- Curated through structured initiatives: Finance Commons and Legal Commons.
- Explicitly includes data from multiple geographic sources: US, France, European, and international organizations.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- The description notes the dataset 'mostly features' data from specific regions, indicating potential geographic bias.
- Row count, file formats, and license information are unknown, which may limit suitability assessment.
Provenance
- Source
- AgentPublic via Hugging Face.
- Collection Method
- Aggregated from open data programs.
- Time Range
- null
- Freshness
- Last updated 2025-01-31 03:54:28; freshness should be verified.
- Geography
- Primarily US, France, European, and international organizations.