A. Kalai, Ofir Nachum, S. Vempala, Edwin Zhang
An incentive-based theoretical analysis showing that accuracy-based evaluations reward guessing in LLMs, leading to hallucinations.
LLMs still produce confident falsehoods (hallucinations), and existing approaches fail to fully explain the cause.
Using learning theory, we prove that next-word prediction training creates unavoidable errors for facts lacking repeated support, and that accuracy metrics reward guessing. We propose 'open rubric' evaluations to incentivize models to admit uncertainty.
By reframing hallucination as an incentive problem, we propose a new evaluation methodology that offers a practical path toward more reliable LLMs.