Results 1 - 10 of 24628
Abstract not available.
Modern AI systems can achieve remarkable predictive and generative performance while remaining unreliable under distribution shift, novel compositions, or intervention. A fundamental reason is non-identifiability: many different internal representations and mechanisms can explain the same observed data and training objective, while implying very different behavior outside the training distribution. Thus, before asking whether a learned representation is interpretable, truthful, or causally meaningful, we should ask a more basic question: what aspects of what the model has learned are actually determined by the available evidence?
I will discuss a line of work that uses identifiability and causality as organizing principles for reliable representation learning. The central idea is that additional structure, such as sparsity, variation across environments, weak supervision, and assumptions about how causal mechanisms change, can turn otherwise underdetermined latent representations into recoverable, semantically meaningful abstractions. I will describe examples ranging from discovering latent causal variables and mechanisms to identifying interpretable concepts inside foundation models, where controlled contrasts and environmental variation can distinguish causally active concepts from merely correlated directions.
This perspective suggests a route from prediction to reliable autonomy: rather than relying solely on increasingly capable models together with post-hoc checks, we can ask what latent variables, concepts, and mechanisms are identifiable; what interventions they support; and which aspects of their behavior can therefore be expected to remain stable as the world changes.
Abstract not available.
Modern machine learning models are increasingly deployed in dynamic environments where reliability must be established continuously rather than assumed. In this talk, I will present two recent advances in sequential statistical inference that address complementary aspects of this challenge. I will first introduce efficient sequential hypothesis tests for continuously auditing deployed machine learning systems, enabling rapid detection of performance degradation while maintaining rigorous Type I error guarantees. I will then show how similar sequential perspectives lead to parameter-free online conformal prediction methods that provide adaptive, group-conditional uncertainty quantification under evolving data distributions. Together, these results illustrate how sequential inference provides a principled statistical framework for moving from detecting when models fail to quantifying when individual predictions can be trusted.
Modern language models are expected to generate outputs that are both valid and diverse, even though they are trained on imperfect data. We study these goals through the framework of language generation in the limit. We first ask whether a generator trained on correct data can cover the full breadth of its target language while avoiding hallucinations. We show that these goals conflict in general: requiring only valid outputs can force the generator to collapse onto a narrow part of the language. We then study generation from data containing incorrect examples and characterize how the density of these errors affects diverse and reliable generation. Finally, since hallucinations may be unavoidable in certain settings, we ask whether they can at least be detected. We show that this problem is tightly connected to the classical setting of language identification in the limit.
This talk is based on joint work with Alkis Kalavasis, Amin Karbasi, Anay Mehrotra, Omar Montasser, John Sous, Xifan Yu, and Felix Zhou.
Language generation in the limit provides an elegant framework for studying the algorithmic problem of generating unseen valid strings from an unknown target language given example strings from that language. In the original formulation of Kleinberg and Mullainathan and nearly all subsequent work, all errors made by the language generator must be confined to a finite prefix of its infinite output. This contrasts with our experience with modern language models which repeatedly hallucinate in practice.
We study the complementary setting in which the language generator may hallucinate infinitely often and investigate how the rate of hallucination affects the class of languages that can be generated. We show that some language collections cannot be generated with finitely many hallucinations but can be generated with infinitely many, even when the hallucinations form a measure-zero subset of the output. More generally, we establish a strict hierarchy of (uncountable) language collections characterized by hallucination rate. This hierarchy also extends to the width of language generation, i.e., the fraction of the target language generated. These results identify hallucination rate as a fundamental parameter in the theory of language generation.
This talk is based on joint work with Fan Wei and Ian Zhang.
Abstract not available.