Sign in to view source links and access this dataset
Description
A large-scale, quality-filtered cybersecurity corpus designed for continual pre-training of large language models. It contains approximately 15.5 billion tokens, including ~4.6B English text, ~5.5B Chinese text, ~4.7B code, and ~0.84B seed tokens. The dataset was created by WhitzardAgent and last updated on June 6, 2026.
Use Cases
Continual pre-training of LLMs based on the described cybersecurity domain knowledge.
Training multilingual cybersecurity assistants based on the English and Chinese text tokens.
Improving code generation and analysis for security tools based on the included code tokens.
Benchmarking model performance on cybersecurity tasks using the curated seed tokens.
Strengths
Contains approximately 15.5 billion total tokens.
Includes distinct token categories: ~4.6B English text, ~5.5B Chinese text, ~4.7B code, and ~0.84B seed tokens.
Sourced from multiple large-scale datasets including Nemotron-CC-v2, Ultra-FineWeb, Fineweb-Edu-Chinese-V2.1, and StarCoderData.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
Source
Extracted from Nemotron-CC-v2, Ultra-FineWeb, Fineweb-Edu-Chinese-V2.1, StarCoderData, and curated seed sources.
Collection Method
Quality-filtered extraction and aggregation from listed sources.
Freshness
Last updated 2026-06-06 09:15:23; freshness should be verified.
License is unknown; terms of use must be verified before application.