Rajveer-code's TrustShift Benchmark provides standardized model predictions for a cross-domain audit of accuracy, calibration, and subgroup reliability under deployment shift. The dataset supports a study finding that the type of distribution shift determines which axis of trustworthiness fails. It was last updated on July 13, 2026.
Use Cases
- Auditing model calibration under deployment shift based on the study's focus on calibration reliability.
- Evaluating subgroup performance disparities based on the description of subgroup reliability analysis.
- Comparing model failure modes across concept, novel-label, and covariate shifts based on the central finding about shift types.
- Benchmarking model robustness using the standardized predictions provided by the benchmark.
Strengths
- Dataset is directly linked to a research study with a defined central finding about distribution shift types.
- Associated code and audit protocol are available via a provided GitHub repository link.
- Last update timestamp (2026-07-13 07:55:22) is explicitly provided.
Limitations
- Description metadata is limited; actual data quality requires manual inspection after download.
- Column-level documentation is absent; field semantics must be inferred after download.
- Row count, file formats, and license information are unknown, which may limit suitability assessment.
Provenance
- Source
- Rajveer-code on Hugging Face
- Collection Method
- Standardized model predictions generated for the TrustShift study.
- Freshness
- Last updated 2026-07-13 07:55:22; freshness should be verified.