Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A benchmark for evaluating the faithfulness of large language models, built on a financial knowledge graph. It contains over 480,000 triplet evaluations sourced from 101 S&P 100 companies, including judge reasoning and full document context. The dataset was curated by Domyn and was last updated on October 6, 2025.
License is CC-BY-NC-4.0, which prohibits commercial use.