Over 40,600 tokens of annotated discourse relations are included from Version 2, with an additional 13,000 tokens annotated in Version 3. The Penn Discourse Treebank annotates discourse relations in the Wall Street Journal section of Treebank-2. Rashmi Prasad led the project, which includes tools for annotation, adjudication, and conversion.
Use Cases
- Train discourse parsing models based on annotated explicit and implicit relations.
- Benchmark natural language understanding systems based on the standardized sense labels.
- Study cross-linguistic discourse patterns based on the corpus's influence on resources in other languages.
- Analyze the frequency and usage of discourse connectives based on the provided statistics.
Strengths
- Annotations are byte-indexed into the raw WSJ text files, providing precise alignment.
- Version 3 includes standardized pairwise annotations, new senses, and consistency checks.
- The corpus includes an annotator tool for viewing and a conversion tool for Version 2 files.
Limitations
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count is unknown, which may limit suitability assessment.
- Data may reflect temporal and domain bias inherent to the Wall Street Journal corpus.
Provenance
- Source
- Penn Discourse Treebank project, based on the Wall Street Journal section of Treebank-2.
- Collection Method
- Manual annotation of discourse relations, followed by adjudication and consistency checks.