Sign in to view source links and access this dataset
Description
1,662,448 articles were harvested from 930 random public MediaWiki instances across the Internet. The collection was created by nyuuzyou, extracting current page content with text, metadata, and structural information. The dataset was last updated on 2026-03-01.
Use Cases
Train language models on diverse, real-world wiki text mentioned in the description.
Analyze knowledge representation and structure across different wiki platforms.
Study multilingual content distribution and topics from the cross-section of public wikis.
Benchmark information retrieval systems on unstructured web-based knowledge sources.
Strengths
Large scale with 1,662,448 articles.
Diverse source base from 930 independent wiki instances.
Includes article text, metadata, and structural information as stated.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is known, but specific data quality, language distribution, and topic balance require manual inspection.
Freshness should be verified; last update is 2026-03-01.
Provenance
Source
930 random public MediaWiki instances across the Internet.
Collection Method
Extraction of current page content from wikis.
Time Range
Current at time of harvest; specific temporal coverage is unknown.