Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Two datasets from different domains designed for multi-hop reading comprehension, where models must combine evidence from multiple documents. The best model accuracy on an annotated test set is 54.5%, compared to human performance at 85.0%. The dataset was created by Johannes Welbl of University College London to investigate the limits of existing text understanding methods.
License is listed as Open Access (green); specific terms should be verified.