Social media text data likely contains examples of toxic language in Indonesian. The dataset is hosted on Kaggle, but its specific size, creation date, and author are unknown. Columns and sample data are unavailable for review.
Use Cases
- Training a toxicity classifier for Indonesian social media posts (inferred from domain, verify after download)
- Analyzing patterns of abusive language in a non-English online context (inferred from domain, verify after download)
- Benchmarking multilingual hate speech detection models (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform for sharing machine learning datasets.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Row count is unknown, which may limit suitability assessment.
- Column-level documentation is absent; field semantics must be inferred after download.
Provenance
- Geography
- Indonesia (inferred from title)