Results 191 - 200 of 24284
Self-attention is the key algorithmic module which powers the Transformer neural architecture. However, the softmax nonlinearity appearing in self-attention makes theoretical analysis challenging. We study the training dynamics of gradient descent in a...
We analyze the landscape and training dynamics of diagonal linear networks in a linear regression task, with the network parameters being perturbed by isotropic normal noise during training. The addition of such noise may be interpreted as a stochastic...
This talk sketches a program extending the concepts of “dominance” and “admissibility” from statistical decision theory to machine learning. I will do so by presenting two instance-wise risk comparison results in linear regression (with Gaussian design)...
No abstract available.