Results 41 - 50 of 24600
No abstract available.
No abstract available.
Large language models are typically evaluated using accuracy-based metrics that reward confident guessing while giving little or no credit for acknowledging uncertainty. Because evaluations influence which models and methods are selected, these incentives can work against techniques that reduce hallucinations at the cost of answering fewer questions. The talk discusses approaches to evaluation that better reward useful uncertainty and reflect the differing costs of mistakes, correct answers, and abstentions, along with open questions in designing evaluations whose incentives better match real-world objectives.
Based on joint work with Santosh Vempala, Ofir Nachum, and Edwin Zhang.
We analyze hallucination in language models from a mathematical perspective and find that the statistical pressures of the next-word prediction training pipeline induce hallucinations under fairly general conditions even when trained on clean data. The core proof is a simple reduction from a natural classification task ("fact or hallucination") to a generation/prompt completion task.
The talk is based on facts and on work with Adam T. Kalai (STOC 2025) and with Kalai, Ofer Nachum and Eddie Zhang (Nature 2026).
LLM agents are increasingly deployed on long-horizon tasks with tool use, irreversible actions, and unpredictable feedback. Yet we have few principled ways to tell, mid-episode, whether an agent is on track or quietly failing. Most uncertainty quantification (UQ) research still centers on single-turn QA, a poor match for interactive agents. In this talk, I'll present a general formulation of agent UQ and the challenges unique to agentic settings, from choosing uncertainty estimators to modeling how uncertainty evolves over an interaction. I'll then show that a powerful answer has been hiding in plain sight: RL post-training already yields an implicit step-level signal, the progress advantage, which recovers the optimal advantage function with no annotation or reward-model training. Across test-time scaling, UQ, and failure attribution, this free byproduct beats confidence baselines and even dedicated trained reward models.