Sign in to view source links and access this dataset
Description
A dataset for binary classification of toxicity in user messages to an online cinema's support chat. Toxic messages were synthetically generated by a local Qwen/Qwen3-4B-Instruct-2507 model across six categories, while non-toxic messages were extracted from MTEB BANKING77 and two Bitext datasets, filtered, and translated into Russian. The dataset was created by aurelianvolturi and last updated on July 12, 2026.
Use Cases
Train a binary classifier to detect toxic messages in Russian support chats.
Benchmark model performance on synthetically generated toxic text across six categories.
Fine-tune language models for content moderation in Russian-language online services.
Strengths
Designed for a specific, practical application: toxicity detection in customer support.
Toxic examples were systematically generated across six distinct categories.
Non-toxic examples were sourced from established datasets and translated into Russian.
Limitations
Row count is unknown, which may limit suitability assessment.
Column-level documentation is absent; field semantics must be inferred after download.
The synthetic generation of toxic messages may not fully capture real-world linguistic patterns.
Provenance
Source
huggingface
Collection Method
Toxic messages synthetically generated by a Qwen model; non-toxic messages sourced from MTEB BANKING77 and Bitext datasets, filtered and translated.
Time Range
null
Freshness
Last updated 2026-07-12 19:38:30; freshness should be verified.
Geography
null
License is unknown; terms of use must be verified before application.