LitBench is a benchmark consisting of 1,980 open-access papers (1,000 epilepsy papers containing answers and 980 decoy papers from unrelated fields) and 9,472 human-reviewed facts. The benchmark includes 2,188 questions requiring information from a single paper and 170 questions requiring information from two papers. Five AI systems were tested: Gemma-4B, Gemma-12B, Sonnet-5, and DeepSeek-V4-Flash. Single-paper accuracy ranged from 55.9% to 91.3% depending on the model, while two-paper accuracy was significantly lower (5.7% to 37.8%). No system was reliable across all tested aspects. The benchmark evaluates four dimensions: stated versus interpretation-requiring facts, the number of competing papers, single-paper versus two-paper evidence, and the ability to refuse answering when evidence is absent.