Sign in to view source links and access this dataset
Description
SovNodeAI's Certified Document QA dataset contains over 6,000 rows of machine-checkable question-answer claims for verifying large language model outputs. Every claim includes a certificate allowing item-by-item re-verification, and the data includes filings newer than major model training cutoffs. The dataset also includes a free 127,000-token verified long-context task set and a frontier failure table comparing six models on 100 questions.
Use Cases
Benchmarking LLM factual accuracy based on span-verified and absence-aware claims.
Training or fine-tuning models for verified document QA using machine-checkable certificates.
Analyzing model failure modes on recent documents using the included frontier failure table.
Developing long-context reasoning tasks based on the 127K-token verified task set.
Strengths
Contains over 6,000 certified rows where every claim is machine-re-checkable.
Includes a free supplementary 127,000-token verified long-context task set.
Data includes filings newer than every major model training cutoff, suggesting recent temporal coverage.
Provides a frontier failure table comparing six models on 100 real questions.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count beyond '6,000+' is unknown, which may limit suitability assessment.
Description metadata is limited; actual data quality and structure require manual inspection after download.
Provenance
Source
SovNodeAI via Hugging Face.
Collection Method
Likely curated for evaluating and verifying language model outputs on documents.
Time Range
Includes filings newer than major model training cutoffs, suggesting recent coverage.
Freshness
Last updated 2026-07-14 15:26:06; freshness should be verified.
License is unknown; terms of use must be verified before application.