Description

Recent results on inference time algorithms have shown an interesting result: the score operators for different distributions can be cast as evaluations of the same general operator at different points in space, we call this a score symmetry. The score of a certain distribution π at a location A is equal to the score of another distribution μ at a location B. If the score of μ has been accurately learned at this equivalent location, we can run the sampling algorithm for π as well. As such, properly learning this underlying general operator allows sampling for not just of a single distribution, but a whole family of distributions defined by linear tilts. This can be applied to reward tilting, allowing novel generation of a small subclass within a training set without memorization. The difficulty with applying this idea in practice is the evaluation points of one density and the training set for another do not overlap. In sampling π I may encounter a location A, and there may be an equivalent location B, but if B does not live in the training set of μ my score estimator will not be accurate at B.


In this work, we show that by restricting ourselves to discrete densities on the hypercube {−1, 1}^d, which is natural for token and language tasks, additional score symmetries arise. Thus there are other locations C, D etc that are also equivalent, and these locations can be organized to live in the training distribution.

 

This seminar is part of the ML and AI Theory series. 

All scheduled dates:

Upcoming

No Upcoming activities yet

Past


Continuous Diffusion on the Hypercube: A Single Tilt–Mean Operator and Training-Free Reward Tilting