An artificially generated dataset for classification tasks related to student dropout. The dataset was created for machine learning practice and is hosted on Kaggle. Specific details regarding its size, creation date, and authorship are not provided in the available metadata.
Use Cases
- Train a binary classifier to predict student dropout based on synthetic features.
- Benchmark different classification algorithms on a controlled, artificial education dataset.
- Simulate and study the impact of various factors on student retention using generated data.
Strengths
- Data is explicitly generated for classification tasks, indicating a clear intended use.
- Being synthetic, the data likely allows for controlled experimentation without privacy concerns.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- Kaggle
- Collection Method
- Artificially generated.