Abstract

Standard paradigms for LLM alignment and evaluation, such as RLHF for alignment and Elo-style rankings for leaderboard evaluation, implicitly assume a single underlying user utility model, an assumption that breaks down under heterogeneous human preferences. As a result, these methods can systematically fail to optimize even for average user satisfaction. In this talk, I will take a social choice perspective to revisit alignment and evaluation under heterogeneous preferences.

I will first discuss pluralistic alignment and introduce the distortion of AI alignment, a framework that captures the worst-case ratio between the optimal achievable average utility and the average utility of the learned policy. This framework characterizes information-theoretic limits of learning from pairwise comparisons, and draws sharp distinctions between alignment methods, showing that Nash Learning from Human Feedback is provably optimal, whereas standard approaches like RLHF and DPO can suffer high or even unbounded distortion. I will then turn to leaderboard-based evaluation and discuss the construction of pluralistic leaderboards, which aim to produce a single global ranking while ensuring fair and stable representation of diverse user populations.

Video Recording