AI & Computing · 2020-05
The Scaling HypothesisDownload sources
Gwern Branwen
Gwern examines the possibility that larger training runs can unlock broader language-model capabilities. The essay collects observations and arguments about what scaling might explain and where it might lead.
Its research-notebook quality is part of the appeal. Keep the evidence, interpretation, and extrapolation separate, and remember that a living essay may differ from its original publication context.
No hosted PDF. Read the original source.