Abstract

Synthetic data such as AI-generated texts and images are increasingly common on the Internet. Recent studies show that if models are iteratively retrained on Internet data mixed with these synthetic data, their performance may deteriorate over time —- a concerning phenomenon often referred to as model collapse.

However, in practice, human engagement with Internet content (such as upvotes, likes, and follow-up interactions) can signal the quality of synthetic data. Can we save generative models from collapse by iteratively retraining them only on such “human-verified” synthetic data? Toward a principled understanding, we investigate this question in the fundamental linear regression setting, showing that retraining on verified synthetic data can avoid model collapse and, in fact, may even yield performance improvements. Our experiments across linear regression, Variational Autoencoders (VAEs) trained on MNIST, and fine-tuning SmolLM2-135M on the XSUM task confirm these theoretical insights.

Video Recording