The Scaling Hypothesis

Gwern Branwen · Independent · 2020-05-28

Read on gwern.net

Written in the immediate aftermath of GPT-3, Gwern's essay is the document that converted "scale matters" from a hunch into a formal hypothesis the field could argue against. He lays out the blessings of scale, contrasts the scaling-pilled and scaling-skeptical camps, and predicts much of what 2021–2025 then delivered.

It is long and unapologetically dense, but no shorter piece gives a learner the same lineage. If you want to understand why OpenAI, Anthropic, DeepMind and xAI all behave the way they do — and why critics like Marcus or Chollet are pushing against this specific argument rather than a strawman — start here.

> Intelligence is "just" simple neural units & learning algorithms applied to diverse experiences at a (currently) unreachable scale.