209,267 entries form this automatically curated dataset for fine-tuning large language models on cybersecurity tasks. It aggregates data from sources including exploitdb, NVD, and MITRE, and is authored by trivedivatsal63. The dataset was last updated on June 11, 2026.
Use Cases
- Fine-tuning LLMs for threat intelligence report generation based on aggregated data from multiple security sources.
- Training models for vulnerability analysis and scoring based on data from NVD and exploitdb.
- Developing automated defensive security tools based on curated data from sources like Cortex and AlienVault.
Strengths
- Contains 209,267 total entries, providing a substantial corpus for model training.
- Aggregates data from 8 distinct sources, including exploitdb (123,508 entries) and NVD (75,684 entries).
- Has a defined version (1.0) and is licensed under Apache 2.0.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count per source is provided, but the specific data structure and features are unknown.
- Freshness should be verified as the last update timestamp is in the future (2026-06-11).
Provenance
- Source
- Aggregated from multiple open sources including exploitdb, NVD, Cortex, MITRE, AlienVault, CTF, OWASP, and Certin.
- Collection Method
- Automatically curated.
- Freshness
- Last updated 2026-06-11 12:02:03