Abstract
Decoupled training methods, such as DiLoCo, relax the requirement for frequent synchronization, allowing large-scale models to be trained across distributed, high-latency compute clusters. This talk examines recent progress in making these methods work at scale and the ongoing development of "optimization for decoupled training." We then explore how the empirical success of decoupled training has unexpectedly influenced "standard" optimization, leading to new lookahead-inspired Nesterov-style methods. Finally, we highlight open questions at the intersection of decoupled training and federated learning. Throughout, we discuss key similarities and distinctions between these fields and ruminate on how heterogeneity, typically seen as a hurdle in federated settings, might actually serve as a powerful tool for improving decoupled training when injected intentionally.