A dataset for binary classification of Russian text toxicity. The toxic class is derived from the AlexSham/Toxic_Russian_Comments dataset, while the non-toxic class is derived from the MTS-AI-SearchSkill/MTSBerquad dataset of neutral user questions. The dataset was authored by molyalya and last updated on 2026-07-15.
Use Cases
- Train binary text classifiers for toxicity detection based on Russian text data.
- Benchmark model performance on Russian toxicity classification tasks.
- Analyze linguistic patterns of toxicity in Russian online comments.
Strengths
- Explicitly designed for binary classification of Russian text toxicity.
- Sources are specified: toxic class from AlexSham/Toxic_Russian_Comments and non-toxic class from MTSBerquad.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Last updated 2026-07-15 11:24:51; freshness should be verified.
Provenance
- Source
- Combined from AlexSham/Toxic_Russian_Comments and MTS-AI-SearchSkill/MTSBerquad datasets.
- Collection Method
- Likely aggregated and labeled from the source datasets.
- Freshness
- 2026-07-15 11:24:51