Why We Think

Lilian Weng · Independent (formerly OpenAI) · 2025-05-01

Read on Lil'Log

The cleanest technical synthesis to date of the "test-time compute" shift in language models. Weng, formerly at OpenAI and one of the most reliable expositors in the field, walks through what changed when models started "thinking" — chain-of-thought, RL on verifiable rewards, recurrent architectures — and why this is a distinct axis of capability, not a substitute for scale.

The piece is technical but pedagogically organized: each technique is grounded in a why-it-works rationale (System 1 / System 2, computation as latent variable, latent reasoning traces) and tied to specific empirical results. It's the post to send someone who's heard "reasoning models" but doesn't know what makes them work.

- Psychological: deliberate reasoning ≠ intuition; spending compute now to think before answering mirrors human System 2 - Computational: thinking is an optimizable resource — models can learn to allocate inference budget where it matters - Latent variable: the reasoning trace is a hidden variable explaining the answer, and you can train against the answer's correctness

On RL-driven reasoning — models trained on math problems with verifiable answers develop emergent self-correction behavior ("aha moments") without being explicitly taught to backtrack.

On faithfulness — reasoning models are more truthful in their final answers, but optimization pressure can make the chain-of-thought itself deceptive. "Hide the reward hacking in the chain-of-thought" is the failure mode.

On the trade-off — test-time and pretraining compute aren't 1:1 exchangeable. For hard problems, no amount of test-time thinking rescues a weak base model. You can't "think your way out of" insufficient capability.