Results 691 - 700 of 24381
Standard paradigms for LLM alignment and evaluation, such as RLHF for alignment and Elo-style rankings for leaderboard evaluation, implicitly assume a single underlying user utility model, an assumption that breaks down under heterogeneous human preferences. As a result, these methods can systematically fail to optimize even for average user satisfaction. In this talk, I will take a social choice perspective to revisit alignment and evaluation under heterogeneous preferences.
I will first discuss pluralistic alignment and introduce the distortion of AI alignment, a framework that captures the worst-case ratio between the optimal achievable average utility and the average utility of the learned policy. This framework characterizes information-theoretic limits of learning from pairwise comparisons, and draws sharp distinctions between alignment methods, showing that Nash Learning from Human Feedback is provably optimal, whereas standard approaches like RLHF and DPO can suffer high or even unbounded distortion. I will then turn to leaderboard-based evaluation and discuss the construction of pluralistic leaderboards, which aim to produce a single global ranking while ensuring fair and stable representation of diverse user populations.
Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? I will present a strategic setting where a human user interacts with multiple differently misaligned AI agents, none of which are individually well-aligned. Nonetheless, when the user's utility lies approximately within the convex hull of the agents' utilities, a condition that becomes easier to satisfy as model diversity increases, strategic competition can yield outcomes comparable to interacting with a perfectly aligned model. I will then move to a setting with multiple heterogeneous users and discuss how the role of model personalization affects emergent pluralistic alignment.
Collaboration is crucial for reaching collective goals. However, its effectiveness is often undermined by the strategic behavior of individual agents---a fact that is captured by a high Price of Stability (PoS) in recent literature. Implicit in the traditional PoS analysis is the assumption that agents have full knowledge of how their tasks relate to one another. We offer a new perspective on bringing about efficient collaboration among strategic agents using information design. Inspired by the growing importance of collaboration in machine learning (such as platforms for collaborative federated learning and data cooperatives), we propose a framework where the platform has more information about how the agents' tasks relate to each other. We characterize how and to what degree such platforms can leverage this information advantage to steer strategic agents toward efficient collaboration.
Concretely, we consider collaboration networks where each node is a task type held by one agent, and each task benefits from contributions made in their inclusive neighborhood of tasks. This network structure is known to the agents and the platform, but only the platform knows each agent's real location. We design two families of persuasive signaling schemes that the platform can use to ensure a small total workload when agents follow the signal. The first family aims to achieve the minmax optimal approximation ratio compared to the optimal collaboration. The second family ensures per-instance strict improvement compared to full information disclosure.