Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Fifteen high-fidelity cardiac surgery scenarios were developed by senior surgeons to benchmark five large language models, including O1 and GPT-4, using a 10-dimensional weighted evaluation framework. Median normalized scores for the top model, O1, reached 0.896, while patient safety and hallucination avoidance were the lowest-scoring dimensions across all models. A separate blinded evaluation by surgeons revealed a 7.57% shift in ratings from affirmative to negative after exposure to expert-curated reference answers.
The primary data files are in DOCX and XLSX formats; the dataset likely contains textual evaluations, scores, and possibly synthetic model outputs rather than raw patient data.