Abstract

Aligning AI systems with human values remains a fundamental challenge, but does our inability to create perfectly aligned models preclude obtaining the benefits of alignment? I will present a strategic setting where a human user interacts with multiple differently misaligned AI agents, none of which are individually well-aligned. Nonetheless, when the user's utility lies approximately within the convex hull of the agents' utilities, a condition that becomes easier to satisfy as model diversity increases, strategic competition can yield outcomes comparable to interacting with a perfectly aligned model. I will then move to a setting with multiple heterogeneous users and discuss how the role of model personalization affects emergent pluralistic alignment.

Video Recording