A federated, non-IID benchmark for phishing URL classification derived from two public Hugging Face datasets. The dataset contains URL strings, binary phishing labels, and a client_id field assigning each example to one of 100 simulated clients. It was created by flwrlabs and last updated on May 11, -2026.
Use Cases
- Benchmarking federated learning algorithms based on the non-IID client assignment.
- Training phishing URL classifiers based on the provided URL strings and binary labels.
- Simulating realistic federated learning scenarios for cybersecurity based on the 100 simulated clients.
Strengths
- Contains a client_id field for simulating federated learning across 100 distinct clients.
- Derived from two established public datasets, suggesting a foundation in prior work.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license information are unknown, limiting suitability assessment.
Provenance
- Source
- flwrlabs on Hugging Face, derived from ealvaradob/phishing-dataset and kmack/Phishing_urls.
- Collection Method
- Federated benchmark created by merging and processing two public datasets.
- Time Range
- null
- Freshness
- Last updated 2026-05-11 10:37:03; freshness should be verified.
- Geography
- null