F1ow421 aggregated and cleaned Russian-language comments from various social platforms and annotated corpora. The dataset is labeled for three independent toxicity classes: profanity, threat, and illegal content. It was last updated on the platform in July 2026.
Use Cases
- Train a multi-label classifier for Russian toxic comment detection based on the three defined categories.
- Benchmark model performance on the specific subtasks of profanity, threat, and illegal content identification.
- Analyze the co-occurrence and distribution of different toxicity types in Russian online discourse.
- Fine-tune a Russian-language BERT model for content moderation tasks.
Strengths
- Data is cleaned of garbage such as links, tags, and extra spaces.
- Features multi-task labeling across three distinct toxicity classes.
Limitations
- Row count, column definitions, and file formats are unknown, limiting suitability assessment.
- The aggregation sources are unspecified, which may introduce source-specific biases.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- Aggregated from various Russian-language comment sources including social platforms and annotated corpora.
- Collection Method
- Aggregated, cleaned, and annotated by author F1ow421.
- Time Range
- null
- Freshness
- Last updated 2026-07-17 17:57:11; freshness should be verified.
- Geography
- Likely focuses on Russian-language content, but specific geographic coverage is unknown.