Sign in to view source links and access this dataset
Description
Over 2 billion public messages were collected from 3,167 distinct public Discord servers, accompanying a paper submitted to ICWSM 2025. The dataset is organized into individual JSON files per server, with an overview file providing server metadata. It was uploaded by author fvdfs41 and last updated on June 9, 2025.
Use Cases
Analyze community structure and evolution based on server metadata and message volume.
Study linguistic patterns and discourse in large-scale online conversations.
Train or benchmark large language models on informal, multi-community text data.
Investigate information diffusion and topic trends across different Discord servers.
Strengths
Contains over 2 billion messages, providing a large-scale text corpus.
Covers 3,167 distinct public Discord servers, offering a multi-community perspective.
Includes a separate metadata file (servers.json) with server descriptions and guides.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
The dataset consists solely of public messages, which may not represent the full activity or private channels of these servers.
Provenance
Source
Public Discord servers.
Collection Method
Data collection method is described in the accompanying ICWSM 2025 paper.
Freshness
Last updated 2025-06-09 23:52:38; freshness should be verified.
License is unknown; users must verify terms of use before downloading.