Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A multi-domain German-English parallel corpus introduced by Aharoni and Goldberg in 2020. The data split was created to avoid duplicate examples and leakage from the train split to the dev/test splits. The original multi-domain data first appeared in Koehn and Knowles (2017) and consists of five datasets from the Opus website.
License is unknown and must be verified before use.