MediNjuzMediNjuz Back to news list

LitBench: A Benchmark for Retrieval-Grounded Multi-Paper Evidence Synthesis in Epilepsy

Source: medRxiv

Original: https://www.medrxiv.org/content/10.64898/2026.09.17.26363288v1?rss=1...

Published: 2026-09-23

LitBench is a benchmark consisting of 1,980 open-access papers (1,000 epilepsy papers containing answers and 980 decoy papers from unrelated fields) and 9,472 human-reviewed facts. The benchmark includes 2,188 questions requiring information from a single paper and 170 questions requiring information from two papers. Five AI systems were tested: Gemma-4B, Gemma-12B, Sonnet-5, and DeepSeek-V4-Flash. Single-paper accuracy ranged from 55.9% to 91.3% depending on the model, while two-paper accuracy was significantly lower (5.7% to 37.8%). No system was reliable across all tested aspects. The benchmark evaluates four dimensions: stated versus interpretation-requiring facts, the number of competing papers, single-paper versus two-paper evidence, and the ability to refuse answering when evidence is absent.