AI & Computing · 2022-03
Training Compute-Optimal Large Language Models (Chinchilla)PDF Download
Hoffmann et al. (DeepMind)
Chinchilla revisits the balance between model size and training data under a fixed compute budget. The central argument is that many large models were trained on too few tokens.
Read it beside Scaling Laws to see scientific progress as a revision of an influential recipe. The practical lesson is to ask how compute is allocated, not only how many parameters a model has.
5.7 MB · SHA-256: 3fd3632a8ef48171bd25282990221d49535d75356192f068b3b2ebe08f2aedd4