Abstract
Recent approaches to AI alignment rely on collecting preference data from humans, then training AI systems to learn preferences and optimize for them. In practice, this is often done by eliciting human choices between pairs of finite-length, or partial, trajectories of actions and observations. Previous theoretical work has studied the problem of learning preferences from choices between partial trajectories, but typically assuming the representation of the human’s state space to be known a priori. This is a limiting assumption; in general, the human may use an internal state representation that is unknown to the agent. Instead, we study a more general model of learning human preferences, where the human maintains an internal representation of the state that is unknown. We prove that it is possible to learn human preferences when the human’s preferences are representable by an (unknown) finite state machine, and we provide a consistent algorithm that recovers the preference relation from choice data. We complement our theoretical results with didactic experiments in simple environments that validate our theory.
Joint work with Karim Abdel Sadek (Berkeley) and Benjamin Van Roy (Stanford).