PAN20 Authorship Analysis: Style Change Detection at Paragraph Level
by Eva Zangerle / Universität Innsbruck
Available on 1 platform
Sign in to view source links and access this dataset
Description
PAN20 Authorship Analysis data supports the detection of author changes in multi-author documents at the paragraph level. The dataset includes two topical subsets, one narrow (technology) and one wide (adding travel, philosophy, economics, history), each split into training, validation, and test sets. The data was provided by Eva Zangerle of Universität Innsbruck for the PAN 2020 shared task.
Use Cases
Train models to detect single- vs. multi-authored documents based on paragraph-level style analysis.
Develop algorithms to identify the exact paragraph boundaries where authorship changes occur.
Benchmark style change detection systems on documents with varying topical breadth as described.
Strengths
Includes ground truth data for training and validation sets, which cover 75% of the total data.
Documents contain zero up to ten style changes from at most three different authors, providing a controlled testbed.
Offers two datasets differing in topical breadth, allowing for analysis of topic influence on style detection.
Limitations
The exact number of rows, columns, and total size of the dataset is unknown.
Column-level documentation is absent; field semantics must be inferred after download.
Last update date is unknown; freshness unverified.
Provenance
Source
Eva Zangerle, Universität Innsbruck, via the PAN 2020 shared task.
Collection Method
Likely compiled for the PAN 2020 shared task on style change detection.
License is listed as Open Access (green); specific terms should be verified upon download.