MediNjuzMediNjuz Back to news list

Evaluating large language models in specialized myopia knowledge and clinical reasoning

Source: Frontiers Medicine

Original: https://www.frontiersin.org/articles/10.3389/fmed.2026.1886498...

Published: 2026-08-06T00:00:00Z

The study evaluated six large language models (ChatGPT, DeepSeek, Gemini, Claude, Qwen, and Doubao) on their ability to provide accurate information and clinical reasoning regarding myopia. ChatGPT achieved the highest accuracy with 92.33% in standardized tasks, 85% in case-based clinical tasks, and 88% in external validation tests. The models showed the strongest performance in core judgment but declined in preferred management decisions. In open-ended case assessments, ChatGPT, DeepSeek, Claude, and Gemini achieved higher expert-rated scores than junior physicians, while Doubao and Qwen showed lower performance. The study concluded that although these models can provide useful information to the public, their responses should serve only as reference material and not as a substitute for actual clinical care.