AI & Computing · 2020-01
Scaling Laws for Neural Language ModelsPDF Download
Kaplan et al. (OpenAI)
Kaplan and colleagues study how language-model loss varies with model size, data, and compute. They turn scaling from an intuition into a set of empirical relationships within their experiments.
The curves are a good starting point for thinking about resource allocation. Their assumptions and measurement range matter; a fitted relationship is not a promise about every future model.
2.4 MB · SHA-256: a41bd7877fd1a6bcbba096b2619618bd2f90e02488f2365644903eb2f7c6a494