01 The Big Picture
For a decade, capability came from one dial: make the model bigger and train it longer. Around 2024 a second dial appeared that anyone with an API key can turn: spend more compute at inference time.
Docs 06–08 covered the mechanics and the wall: prefill reads your prompt, decode writes tokens one at a time, and decode is throttled by memory bandwidth. This doc adds the newest axis on top of that machinery. A reasoning model (o1/R1-style) doesn't answer directly — it first emits thousands of thinking tokens: planning, trying approaches, checking its own work, backtracking. None of that changes the weights. It's the same transformer from doc 06, run for more decode steps.
The punchline of the scaling-law literature: for hard problems, an extra 10× of test-time compute can be worth several× of parameter count — and test-time compute is rented by the token, while parameters are owned by whoever trained the model.
02 What It Is (and Is Not)
Test-time compute scaling is the deliberate increase of inference-time computation — more tokens, more samples, more search — to improve answer quality, with frozen weights. Reasoning models are the trained version of this: post-trained with reinforcement learning on verifiable rewards (RLVR) to produce long chains of thought before answering.
What it is not:
Not new parameters
Thinking tokens are ordinary decode steps from the same model (doc 06). "Reasoning" lives in the token stream, not in a new architecture. The model got better at using its own output as working memory — RL taught it which intermediate tokens help.
Not just longer prompts
Asking a standard model "think step by step" gets you a few hundred tokens of post-hoc rationalization. An RL-trained reasoner explores: it can abandon an approach mid-stream, catch its own error, and restart — behavior shaped by reward, not by prompt phrasing.
Three mechanisms, one idea — pay tokens for accuracy:
03 Why Scale at Test Time — the Practical Asymmetry
Why did the field pivot here instead of just training bigger models? Because of an economic asymmetry:
The compute-optimal trade-off. Total performance is a function of parameters N and test-time compute — the frontier is a set of iso-accuracy curves: a small model thinking 50× longer can match a big model thinking 1×. Empirically, accuracy gained per unit of test-time compute follows a power law with a shrinking exponent:
04 How It Works — One Problem, Seven Steps
Step through a reasoning model solving a single hard problem, from prompt to final answer.
Two details matter for engineers. First, the verify step is trained, not prompted — RL on verifiable rewards teaches the model that catching its own errors before committing earns reward, so self-correction emerges at inference. Second, the loop can run many turns, which is why reasoning-model responses take seconds to minutes: it's serial decode, the memory-bound regime of doc 08, for tens of thousands of tokens.
05 The Math of Sampling Several Times
When one sample isn't enough, you pay for k. Here is what each extra sample actually buys.
Pass@k — how good is one model, honestly? Before buying samples you need the per-sample accuracy p. Estimating "probability that k draws contain ≥1 correct" is biased if you measure naive frequencies; the standard unbiased estimator (Codex evaluation, Chao Chen et al.) draws n ≥ k samples and counts c correct ones:
The combinatorial form is exact: each of the C(n,k) subsets is equally likely to be your k draws, and the fraction with zero correct samples is C(n−c,k)/C(n,k). Doc 13 covers why this matters: a model with pass@1 = 20% and pass@50 = 80% is a sampling problem, not a capability problem — and test-time compute is exactly the tool for sampling problems.
Best-of-n with a verifier. Sample n candidates; a verifier (checker code, reward model, or the model's own judge) picks one. If a sample is correct with probability p and the verifier selects a correct candidate when one exists with reliability q:
Because draws are independent but p varies per problem, the exact distribution of the number correct is a Poisson-binomial, not a plain binomial — averaging over problem difficulty. The intuition that carries: P(max of n draws) saturates exponentially — accuracy climbs fast for small n, then flattens. If p ≈ 0, no n saves you; best-of-n multiplies what's there, it doesn't create it.
Majority voting (self-consistency). No verifier — sample n answers, take the majority. If samples are independent-ish, each correct with probability p, the vote is correct whenever the binomial tail exceeds half:
Majority voting wins when errors are diverse (different failure modes per sample — reasoning paths disagree) and p > ½-ish; it fails when errors are correlated (the model systematically misreads the problem — every sample makes the same mistake, and the vote ratifies it). Diversity across samples is why doc 14's temperature matters here: temperature 0 gives you n identical votes.
06 Reasoning-Model KV & Context Mechanics
A 30K-token thinking chain is 30K KV cache entries you pay for on every subsequent decode step.
Doc 06 established the anatomy: every generated token attends over the full KV cache, so decode cost per step grows linearly with context length — and doc 08 showed the binding constraint is memory bandwidth, not FLOPs. Reasoning models push all three levers at once:
07 When It Helps, When It Doesn't — and the Economics
Why RLVR beats naive RLHF for reasoning. Training long chains of thought needs a reward that can't be gamed. Verifiable reward r = 1[answer correct] — unit tests, exact answers, formal proofs — ties reward to reality. A learned human-preference reward model instead scores plausibility, and a long confident-sounding chain optimizes plausibility: reward hacking. RLVR's reward is cheap, exact on its domain, and unhackable within it — at the cost of working only where checkers exist. That boundary is precisely the table below.
| Domain | Verifier exists? | Test-time compute payoff |
|---|---|---|
| Math / proofs | Exact (checker, formal system) | Large — AIME-style jumps of 2–4× from thinking budgets |
| Code | Exact (tests, compilation) | Large — self-generated tests let the model verify itself |
| Science / analysis | Partial (units, consistency checks) | Moderate — helps structured reasoning, can't check facts |
| Summarization / NER / chat | No — reward is fuzzy human preference | Near zero or negative — empirically overthinking: accuracy flat, cost 10–50×, sometimes worse |
The economics. Thinking tokens are billed as output tokens — the 3–5× row of doc 07's table. The number that actually governs your purchase decision is cost per correct answer:
Worked example: baseline gives 60% correct at 400 output tokens; a reasoning model gives 92% at 6,000 tokens. Baseline: 400/0.60 ≈ 667 tokens per correct answer. Reasoning: 6,000/0.92 ≈ 6,522. If a correct answer is worth more than ~10× a wrong one (an escalation, a human retry), reasoning wins; if wrong answers are cheap to catch and retry, it doesn't. Extra thinking is not worth it when: p is already high (nothing to gain), the task has no verifier signal to steer the chain, errors are systematic rather than stochastic (majority-vote logic fails), or latency is the product.
Spend thinking tokens on verifiable, multi-step problems; expose "reasoning effort" as a per-task config; measure pass@1 and cost-per-correct, not vibes; batch parallel samples when the checker is cheap.
Route extraction, classification, or summarization through a reasoner; pay for 30K thinking tokens on a retry-able 100-token task; assume longer thinking = better (the α < 1 curve flattens).
08 Mental Models
The student (frozen brain — frozen weights) gets extra scratch paper and more clock time instead of a smarter brain. Hard problems improve a lot; "what is 7×8" doesn't. Lets you reason about: why the payoff concentrates in multi-step verifiable domains, and why easy tasks just get slower and pricier.
Parameters are a factory you must build (capex, months, gating by capital); test-time tokens are a workshop you rent by the hour (opex, instant, per-project). Lets you reason about: the asymmetry of section 03, and why budget flexibility — 10× spend on the one hard problem — often beats a uniformly bigger factory.
09 Common Misconceptions
"Reasoning models are a different architecture." No — same transformer, same decode loop (doc 06). What changed is post-training: RL on verifiable rewards taught the model to use its own tokens as search. "Reasoning" is a behavior, not a module.
"More thinking tokens always mean better answers." The scaling curve J(c_test) ∝ c_test^α has α < 1 and flattens; beyond a problem-dependent point you pay 2× tokens for +1% accuracy. And on non-verifiable tasks, overthinking is empirically negative: more room to rationalize a wrong reading.
"Best-of-n / majority voting is obsolete now that models reason." It composes. Reasoning models raise per-sample p; sampling multiplies it — 1−(1−p)ⁿ still applies, just starting from a higher p. Parallel sampling is also latency-flat (batch), while serial thinking is not.
"Thinking tokens are free because they're hidden." They're billed as output tokens at full rate, they occupy KV cache for the request's lifetime (section 06), and they add latency linearly. Hidden from your eyes is not hidden from your invoice.
"Self-correction means I can trust the answer." Verification is only as good as the model's ability to check the domain. On verifiable tasks it's strong (tests, arithmetic); on factual recall, a model "checking itself" often re-confirms the same hallucination with more confidence.