Sign in to view source links and access this dataset
Description
The German Innsbruck Corpus (GermInnC) 1800-1950 is a digitized corpus of approximately 840,000 tokens of German text. It was created by Konstantin Niehaus at the University of Innsbruck and is balanced by period, region, and genre. The corpus includes seven genres such as drama, newspapers, and scientific texts, covering three periods and five German-speaking regions.
Use Cases
Study historical language variation based on the corpus's balance across five German-speaking regions.
Analyze genre-specific language change based on the seven text categories like drama, newspapers, and sermons.
Train models for historical German NLP tasks based on the lemmatized and linguistically annotated text versions.
Research standardisation processes in German based on the corpus's coverage from 1800 to 1950.
Strengths
Approximately 840,000 tokens of text.
Balanced design across three periods (1800-1850, 1851-1900, 1901-1950), five regions, and seven genres.
Available in raw, lemmatized, and fully annotated versions with linguistic annotation using the Stuttgart Tag Set.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count is unknown, which may limit suitability assessment for some tasks.
Provenance
Source
Universität Innsbruck
Collection Method
Digitized corpus built after the fashion of the German Manchester Corpus (GerManC).
Time Range
1800-1950
Geography
Five German-speaking regions: North German, West Central German, East Central German, West Upper German (including Switzerland), East Upper German (including Austria).
License is listed as Open Access (green); specific terms should be verified.