Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
Fifteen high-fidelity cardiac surgery scenarios were used to evaluate five large language models via a blinded, two-phase framework involving senior surgeons. Median normalized scores across models ranged from 0.521 to 0.896, with scenario comprehension scoring highest and patient safety scoring lowest. The evaluation revealed a judgment shift, with 7.57% of ratings revised from affirmative to negative after surgeons reviewed expert reference answers.
License is CC-BY-4.0. Primary data files are in DOCX and XLSX formats.