MediNjuzMediNjuz Back to news list

Assessing the Performance of Artificial Intelligence on Anesthesiology In-Training Examinations and Applicability in Medical Education

Source: medRxiv

Original: https://www.medrxiv.org/content/10.64898/2026.09.21.26363591v1?rss=1...

Published: 2026-09-22

The study compared the performance of two current-generation large language models (Claude Opus 5 and GPT-5.6) on 1,001 anesthesiology examination questions. Claude Opus 5 correctly answered 947 questions (94.6%) and GPT-5.6 answered 946 questions (94.5%), with no statistically significant difference. Both models achieved approximately 95% accuracy, placing them in the highest category of the examination normative framework, corresponding to approximately the 99th percentile. When figures were provided, accuracy on figure-dependent questions was 94.4% for Claude and 83.3% for GPT-5.6, but without figures, accuracy dropped to 52.8%. The models agreed on 957 of 1,001 questions (95.6%). The findings suggest that current-generation language models have potential to serve as accessible educational tools for anesthesiology trainees, provided complete source materials are available.