A multimodal dataset concerning cyberbullying, likely containing text and other media types from Indonesian-language social media sources. It was published on Kaggle, but the author, organization, and specific collection details are unknown. The last update date, row count, and file formats are also unspecified.
Use Cases
- Train a multimodal classifier to detect cyberbullying in Indonesian social media posts (inferred from domain, verify after download)
- Analyze the relationship between text and image content in abusive online interactions (inferred from domain, verify after download)
- Benchmark models for cross-modal toxicity detection in a low-resource language context (inferred from domain, verify after download)
Strengths
- Published on Kaggle, a platform with established data sharing practices.
- The title indicates a focus on Indonesian-language content, which may address a gap in resources for this language.
Limitations
- Metadata is minimal; actual content requires verification after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
Provenance
- Source
- Kaggle
- Geography
- Indonesia (inferred from title)