Sign in to view source links and access this dataset
Description
2,621,468 question-answer examples for training large language models on cybersecurity topics. The dataset aggregates records from the NIST NVD CVE database and a detailed CVE analysis dataset, covering vulnerabilities, attack techniques, and defensive strategies. It was created by user 'rezaduty' and last updated on Hugging Face in June 2026.
Use Cases
Fine-tuning LLMs for cybersecurity question-answering based on the described Q&A structure.
Training models to analyze Common Vulnerabilities and Exposures (CVEs) based on the inclusion of NVD records.
Developing systems for automated vulnerability detection and remediation suggestion based on the coverage of attack techniques and defensive strategies.
Strengths
Contains 2,621,468 examples, providing a large-scale resource for model training.
Sources data from the authoritative NIST NVD CVE database, covering records from 2002 to 2025.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count and file format details are unknown, which may limit suitability assessment.
Provenance
Source
Primarily the NIST National Vulnerability Database (NVD) CVE records and the AlicanKiraz0/All-CVE-Records-Training-Dataset.
Collection Method
Likely aggregated and formatted for question-answering tasks from existing CVE databases.
Time Range
CVE records from 2002 to 2025.
Freshness
Last updated 2026-06-04 03:09:24; freshness should be verified.
License is unknown; terms of use must be verified before application.