Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
The Artha Personal-Finance Reasoning Benchmark is a reproducible, model-agnostic benchmark for evaluating LLM-based personal-finance agents. It assesses answer quality and grounding/hallucination resistance over a user's full transaction ledger, using a four-dimension rubric scored by a three-model-family LLM judge panel. The dataset was created by Tej-Katika and last updated on Hugging Face on 2026-07-02.
License is unknown; terms of use must be verified before application.