Phishing and Benign Dataset with 210,000 Total Samples
Available on 1 platform
Sign in to view source links and access this dataset
Description
210,000 total samples, including 200,000 benign and 10,000 phishing instances, are available for security model training. The dataset is hosted on Kaggle, but its author, creation method, and specific features are not detailed. Column-level documentation is absent, requiring users to infer the data structure after acquisition.
Use Cases
Training binary classifiers to distinguish phishing from benign content based on the described class labels.
Fine-tuning pre-trained security models using the provided phishing and benign samples.
Evaluating model performance on a dataset with a significant class imbalance favoring benign examples.
Benchmarking feature engineering techniques for phishing detection tasks.
Strengths
Contains 210,000 total labeled instances, providing substantial data volume.
Explicitly includes 10,000 phishing samples, offering a dedicated set of malicious examples.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for the dataset is unknown, which may limit suitability assessment.
Data may reflect temporal or source bias inherent to its collection on Kaggle.
Provenance
Source
Kaggle
Collection Method
Unknown
Time Range
Unknown
Freshness
Last update date is unknown; freshness unverified.
Geography
Unknown
License is unknown; usage restrictions must be verified.