Sign in to view source links and access this dataset
Description
7,447,646 deduplicated documents link hacker community discussions, exploit databases, vulnerability advisories, and fix commits through a shared CVE identifier space. The dataset, created by DatasetSubmission, integrates 64 public sources across 8 layers and includes 360,004 CVE-linked rows. It was last updated on Hugging Face in May 2026.
Use Cases
Predict exploit emergence based on hacker forum discourse mentioned in the description
Analyze the timeline from vulnerability disclosure to fix commit using the linked CVE data
Train models to classify threat severity by correlating advisories with community discussions
Study the propagation of vulnerability information across different source layers
Strengths
Large scale with 7,447,646 exact-deduplicated documents
Multi-source integration from 64 public forum and source identifiers
Long temporal span covering data from 1988 to 2026
Direct linkage of 360,004 rows to Common Vulnerabilities and Exposures (CVE) identifiers
Limitations
Column-level documentation is absent; field semantics must be inferred after download
Row count for the primary data table is unknown, which may limit suitability assessment
Data may reflect source bias inherent to the 64 integrated public forums and databases
Provenance
Source
DatasetSubmission on Hugging Face
Collection Method
Likely aggregated from 64 public forums, exploit databases, vulnerability advisories, and code repositories.
Time Range
1988 to 2026
Freshness
Last updated 2026-05-06 21:16:23
License is unknown; terms of use must be verified before application.