IndoDiscourse is a multi-labeled Indonesian text dataset examining toxicity, polarization, and demographic information. The dataset was restructured and expanded by author Exqrch on October 31, 2024, and last updated on the Hugging Face platform in June 2025. It groups unique texts together and includes annotations from multiple annotators.
Use Cases
- Train toxicity detection models based on the toxicity labels mentioned in the description.
- Study polarization in online discourse based on the described 'Polarized' column.
- Analyze the relationship between demographic information and discourse features as suggested by the dataset's scope.
- Benchmark multi-label classification models for Indonesian social text.
Strengths
- Dataset was restructured and expanded with new data on October 31, 2024, indicating active maintenance.
- Texts are grouped uniquely and include annotations from multiple annotators, suggesting a structured annotation process.
- Includes multiple analysis dimensions: toxicity, polarization, and demographics as per the description.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Description metadata is limited; actual data quality requires manual inspection after download.
Provenance
- Source
- huggingface
- Freshness
- Last updated 2025-06-26 07:49:49; freshness should be verified.
- Geography
- Indonesia