Image
Nearly all scaling laws research treats English as the evaluation objective. In this work, we redefine scaling laws to support multilingual objectives, and to inform how practitioners should mix hundreds of language sources to optimize for a non-English language target. Our state-of-the-art scaling laws are also used to determine if its better to pretrain from scratch or finetune from a multilingual checkpoint, and how many languages different model sizes can easily support. We release our cross-lingual transfer matrix for practitioner use and downstream analysis.