Sign in to view source links and access this dataset
Description
A Russian-language dataset for training a multi-task toxicity classifier across three independent binary categories: profanity, threats, and illegal requests. The dataset, created by AtesiT, was last updated on July 15, 2026. It is designed for a multi-label classification task where a single text can belong to multiple categories simultaneously.
Use Cases
Training multi-label classifiers to detect profanity in Russian text based on the described 'profanity' category.
Developing threat detection models for user safety based on the described 'threat' category.
Identifying texts related to illegal activities for automated filtering based on the described 'illegal' category.
Benchmarking the performance of toxicity detection models on a multi-task Russian-language dataset.
Strengths
Defines three specific, independent binary categories for toxicity classification: profanity, threats, and illegal content.
Designed for a multi-label task, reflecting real-world complexity where texts can belong to multiple categories.
Limitations
Row count, column definitions, and sample data are unknown, limiting suitability assessment.
Column-level documentation is absent; field semantics must be inferred after download.
Last updated 2026-07-15 20:18:45; freshness should be verified.
Provenance
Source
huggingface, author AtesiT
Freshness
Last updated 2026-07-15 20:18:45.
Geography
Russian-language texts
License is unknown; terms of use must be verified before application.