MediNjuzMediNjuz Back to news list

The Unreliable Judges: Assessing Reproducibility and Self-Preference Bias of LLMs as Free-Text Evaluators

Source: medRxiv

Original: https://www.medrxiv.org/content/10.64898/2026.06.15.26355670v1?rss=1...

Published: 2026-06-17

The study compares 71 human experts with six large language models (LLMs) in evaluating text responses. Findings show that AI evaluators have a strong tendency to prefer responses generated by AI models. Neither human nor AI evaluations reliably identified whether a response was created by a human or AI. AI scores were strongly correlated with surface-level features such as text length and lexical diversity, whereas human scores did not show this correlation. Analysis of the hidden states of AI models revealed that verbosity is the main factor driving bias. When questions and answers were randomly paired, long responses maintained high scores even when they no longer answered the question, while short responses did not. The research also highlighted that API-based and batch processing increase result unpredictability.