Self-attention is the key algorithmic module which powers the Transformer neural architecture. However, the softmax nonlinearity appearing in self-attention makes theoretical analysis challenging. We study the training dynamics of gradient descent in a softmax self-attention layer trained to perform linear regression and propose a simple first-order optimization algorithm which converges to the globally optimal self-attention parameters at a geometric rate. Our analysis proceeds in two steps. First, we show that in the infinite-data limit the regression problem solved by the self-attention layer is equivalent to a nonconvex matrix factorization problem. Second, we exploit this connection to design Hera, a novel ``structure-aware" variant of gradient descent which efficiently optimizes the original finite-data regression objective. We present simulation results which show that Hera can achieve order-of-magnitude speedups over Muon and Adam on various regression and time-series prediction tasks. Joint work with Mahdi Soltanolkotabi and Peter Bartlett.
This seminar is part of the ML and AI Theory series.
All scheduled dates:
Upcoming
Past
No Past activities yet