Sign in to view source links and access this dataset
Description
Companion records for the paper Token-Set Choice Confounds POPE: A Systematic Audit of Yes/No Extraction in VLM Hallucination Evaluation (Jayakumar & Thilak, 2026). The dataset hosts 9,000 per-question prediction records, diagnostics, ablations, and cross-model audits that back every numeric claim in the paper. Authored by kesav2k04, it was last updated on June 14, 2026.
Use Cases
Auditing the POPE evaluation method based on the described per-question prediction records.
Reproducing numeric claims from the associated paper based on the linked JSON artifacts.
Conducting cross-model comparisons based on the described cross-model audit data.
Analyzing the impact of token-set choice on VLM evaluation based on the described diagnostics and ablations.
Strengths
Contains 9,000 per-question prediction records.
Directly supports every numeric claim in the associated 2026 research paper.
Designed for full reproducibility without re-running experiments.
Limitations
Column-level documentation is absent; field semantics must be inferred after download.
Row count beyond the 9,000 per-question records is unknown, which may limit suitability assessment.
Provenance
Source
huggingface
Collection Method
Created as companion data for a research paper; likely contains experimental results from model evaluations.
Freshness
Last updated 2026-06-14 19:44:55.
License is unknown; terms of use should be verified before application.