Results 11 - 20 of 24633
Abstract not available.
Modern AI systems can achieve remarkable predictive and generative performance while remaining unreliable under distribution shift, novel compositions, or intervention. A fundamental reason is non-identifiability: many different internal representations and mechanisms can explain the same observed data and training objective, while implying very different behavior outside the training distribution. Thus, before asking whether a learned representation is interpretable, truthful, or causally meaningful, we should ask a more basic question: what aspects of what the model has learned are actually determined by the available evidence?
I will discuss a line of work that uses identifiability and causality as organizing principles for reliable representation learning. The central idea is that additional structure, such as sparsity, variation across environments, weak supervision, and assumptions about how causal mechanisms change, can turn otherwise underdetermined latent representations into recoverable, semantically meaningful abstractions. I will describe examples ranging from discovering latent causal variables and mechanisms to identifying interpretable concepts inside foundation models, where controlled contrasts and environmental variation can distinguish causally active concepts from merely correlated directions.
This perspective suggests a route from prediction to reliable autonomy: rather than relying solely on increasingly capable models together with post-hoc checks, we can ask what latent variables, concepts, and mechanisms are identifiable; what interventions they support; and which aspects of their behavior can therefore be expected to remain stable as the world changes.
Modern language models are expected to generate outputs that are both valid and diverse, even though they are trained on imperfect data. We study these goals through the framework of language generation in the limit. We first ask whether a generator trained on correct data can cover the full breadth of its target language while avoiding hallucinations. We show that these goals conflict in general: requiring only valid outputs can force the generator to collapse onto a narrow part of the language. We then study generation from data containing incorrect examples and characterize how the density of these errors affects diverse and reliable generation. Finally, since hallucinations may be unavoidable in certain settings, we ask whether they can at least be detected. We show that this problem is tightly connected to the classical setting of language identification in the limit.
This talk is based on joint work with Alkis Kalavasis, Amin Karbasi, Anay Mehrotra, Omar Montasser, John Sous, Xifan Yu, and Felix Zhou.
Language generation in the limit provides an elegant framework for studying the algorithmic problem of generating unseen valid strings from an unknown target language given example strings from that language. In the original formulation of Kleinberg and Mullainathan and nearly all subsequent work, all errors made by the language generator must be confined to a finite prefix of its infinite output. This contrasts with our experience with modern language models which repeatedly hallucinate in practice.
We study the complementary setting in which the language generator may hallucinate infinitely often and investigate how the rate of hallucination affects the class of languages that can be generated. We show that some language collections cannot be generated with finitely many hallucinations but can be generated with infinitely many, even when the hallucinations form a measure-zero subset of the output. More generally, we establish a strict hierarchy of (uncountable) language collections characterized by hallucination rate. This hierarchy also extends to the width of language generation, i.e., the fraction of the target language generated. These results identify hallucination rate as a fundamental parameter in the theory of language generation.
This talk is based on joint work with Fan Wei and Ian Zhang.
Large language models are typically evaluated using accuracy-based metrics that reward confident guessing while giving little or no credit for acknowledging uncertainty. Because evaluations influence which models and methods are selected, these incentives can work against techniques that reduce hallucinations at the cost of answering fewer questions. The talk discusses approaches to evaluation that better reward useful uncertainty and reflect the differing costs of mistakes, correct answers, and abstentions, along with open questions in designing evaluations whose incentives better match real-world objectives.
Based on joint work with Santosh Vempala, Ofir Nachum, and Edwin Zhang.
Kleinberg and Mullainathan recently revisited a classical model for language identification in the limit, with a modern, language generation lens. They introduced a model for language generation in the limit, an abstraction of modern large language models that has proved to be an interesting theoretical sandbox to investigate mathematical properties of LLMs. This talk will serve as an introduction to this emerging area of research, the first of a three-talk mini-survey on the subject.
We analyze hallucination in language models from a mathematical perspective and find that the statistical pressures of the next-word prediction training pipeline induce hallucinations under fairly general conditions even when trained on clean data. The core proof is a simple reduction from a natural classification task ("fact or hallucination") to a generation/prompt completion task.
The talk is based on facts and on work with Adam T. Kalai (STOC 2025) and with Kalai, Ofer Nachum and Eddie Zhang (Nature 2026).