Sign in to view source links and access this dataset
Description
1,519 questions comprise this benchmark for evaluating long-context, multimodal, and cross-document understanding. The dataset, created by 'anonymous12123' and last updated in May 2026, includes benchmark questions, answers, public source URLs for documents, and human-annotated evidence pages and snippets. It contains 331 public document records in a separate file.
Use Cases
Benchmarking model performance on long-context question answering based on the 1,519 provided questions.
Evaluating multimodal document understanding capabilities using the human-annotated evidence pages and snippets.
Testing cross-document reasoning and information synthesis across the 331 provided public documents.
Training retrieval-augmented generation (RAG) systems using the provided document URLs and evidence annotations.
Strengths
Includes 1,519 benchmark questions specifically for long-context and multimodal understanding.
Provides human-annotated evidence pages and snippets for evaluation.
Contains 331 public document records with source URLs and metadata.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count for the primary QA file is unknown, which may limit suitability assessment.
The author and organization are listed as 'anonymous12123' and unknown, respectively, limiting provenance clarity.
Provenance
Source
huggingface
Collection Method
Likely manually curated and annotated for benchmarking purposes.
Time Range
null
Freshness
Last updated 2026-05-07 11:19:43; freshness should be verified.
Geography
null
License is unknown; terms of use must be verified before application.