AI Learning Series · Part 18

Test-Time Compute & Reasoning Models

The third scaling axis: instead of growing parameters or training data, grow compute at inference. Thinking tokens are decoding decisions — billed as output tokens — that buy accuracy with latency and memory.

Transformers
→
Post-Training
→
Test-Time Compute
→
Long-Context
→
Memory Systems

01 The Big Picture

For a decade, capability came from one dial: make the model bigger and train it longer. Around 2024 a second dial appeared that anyone with an API key can turn: spend more compute at inference time.

Docs 06–08 covered the mechanics and the wall: prefill reads your prompt, decode writes tokens one at a time, and decode is throttled by memory bandwidth. This doc adds the newest axis on top of that machinery. A reasoning model (o1/R1-style) doesn't answer directly — it first emits thousands of thinking tokens: planning, trying approaches, checking its own work, backtracking. None of that changes the weights. It's the same transformer from doc 06, run for more decode steps.

The punchline of the scaling-law literature: for hard problems, an extra 10× of test-time compute can be worth several× of parameter count — and test-time compute is rented by the token, while parameters are owned by whoever trained the model.

02 What It Is (and Is Not)

Test-time compute scaling is the deliberate increase of inference-time computation — more tokens, more samples, more search — to improve answer quality, with frozen weights. Reasoning models are the trained version of this: post-trained with reinforcement learning on verifiable rewards (RLVR) to produce long chains of thought before answering.

What it is not:

Not new parameters

Thinking tokens are ordinary decode steps from the same model (doc 06). "Reasoning" lives in the token stream, not in a new architecture. The model got better at using its own output as working memory — RL taught it which intermediate tokens help.

Not just longer prompts

Asking a standard model "think step by step" gets you a few hundred tokens of post-hoc rationalization. An RL-trained reasoner explores: it can abandon an approach mid-stream, catch its own error, and restart — behavior shaped by reward, not by prompt phrasing.

Three mechanisms, one idea — pay tokens for accuracy:

Serial: long chains of thought. One sample, many thinking tokens. Depth of exploration per attempt — this is what o1/R1-style training buys.
Parallel: sample several times. Best-of-n with a verifier, or majority voting (self-consistency). More attempts, pick better. Doc 14's sampling controls (temperature, top-p) are the knobs here.
Adaptive: budget forcing. Allocate thinking length per problem difficulty — a small question gets 200 tokens, a proof gets 20K. The "reasoning effort" setting on modern APIs is this dial, exposed to you.

03 Why Scale at Test Time — the Practical Asymmetry

Why did the field pivot here instead of just training bigger models? Because of an economic asymmetry:

Anyone can rent tokens; only labs can grow parameters. A frontier training run costs hundreds of millions and is gated by data, chips, and capital. Extra inference compute costs dollars per problem. Test-time scaling is the first capability axis you control directly.
Per-problem budgets beat average budgets. A bigger model spends its extra capacity on every query, easy or hard. Test-time compute lets you concentrate spend: 100 tokens on "summarize this", 30,000 on this one hard integration. Utilization of compute is far better.
It compounds with verifiable domains. Math and code have checkers. A model that can check itself in-context converts raw token spend into genuinely better answers — which is what the RLVR training in the next section exploits.

The compute-optimal trade-off. Total performance is a function of parameters N and test-time compute — the frontier is a set of iso-accuracy curves: a small model thinking 50× longer can match a big model thinking 1×. Empirically, accuracy gained per unit of test-time compute follows a power law with a shrinking exponent:

J(c_test) ∝ c_test^α, α < 1 and falling with difficulty → the first 10× of thinking buys a lot; the next 10× buys less; never zero, but diminishing.
⚖️
Engineering read: test-time compute is a substitute for model size with exchange-rate decay. Cheap substitution at the low end, expensive at the high end — the same diminishing-returns shape as every other scaling law you know.

04 How It Works — One Problem, Seven Steps

Step through a reasoning model solving a single hard problem, from prompt to final answer.

PROMPT problem statement + constraints PLAN decompose → subgoals THINKING TOKENS — long decode, same weights "try approach A… edge case fails… switch to B… recall lemma…" 1K–30K tokens · hidden from you · billed as output DRAFT ANSWER candidate + working shown VERIFY re-derive / check edge cases / unit-test error found → self-correct, loop back FINAL ANSWER verification passed → emit result Total: 1 prompt read (prefill) + 2,000–30,000 decode steps + 0 parameter updates The thinking block is discarded from the visible reply — but every one of its tokens lives in the KV cache until the request ends.

Two details matter for engineers. First, the verify step is trained, not prompted — RL on verifiable rewards teaches the model that catching its own errors before committing earns reward, so self-correction emerges at inference. Second, the loop can run many turns, which is why reasoning-model responses take seconds to minutes: it's serial decode, the memory-bound regime of doc 08, for tens of thousands of tokens.

05 The Math of Sampling Several Times

When one sample isn't enough, you pay for k. Here is what each extra sample actually buys.

Pass@k — how good is one model, honestly? Before buying samples you need the per-sample accuracy p. Estimating "probability that k draws contain ≥1 correct" is biased if you measure naive frequencies; the standard unbiased estimator (Codex evaluation, Chao Chen et al.) draws n ≥ k samples and counts c correct ones:

pass@k = E[1 − C(n−c, k) / C(n, k)] = 1 − Πᵢ₌₀ᵏ⁻¹ (n − c − i)/(n − i) ← evaluates in O(k), no factorials γ_k = Πᵢ₌₀ᵏ⁻¹ (1 − i/n) · [correction for c of n correct]

The combinatorial form is exact: each of the C(n,k) subsets is equally likely to be your k draws, and the fraction with zero correct samples is C(n−c,k)/C(n,k). Doc 13 covers why this matters: a model with pass@1 = 20% and pass@50 = 80% is a sampling problem, not a capability problem — and test-time compute is exactly the tool for sampling problems.

Best-of-n with a verifier. Sample n candidates; a verifier (checker code, reward model, or the model's own judge) picks one. If a sample is correct with probability p and the verifier selects a correct candidate when one exists with reliability q:

P(at least one correct in n) = 1 − (1 − p)ⁿ A(n) ≈ 1 − (1 − p·q)ⁿ ← effective per-sample hit rate p·q p·q > p ⟺ verifier adds signal — otherwise best-of-n hurts

Because draws are independent but p varies per problem, the exact distribution of the number correct is a Poisson-binomial, not a plain binomial — averaging over problem difficulty. The intuition that carries: P(max of n draws) saturates exponentially — accuracy climbs fast for small n, then flattens. If p ≈ 0, no n saves you; best-of-n multiplies what's there, it doesn't create it.

Majority voting (self-consistency). No verifier — sample n answers, take the majority. If samples are independent-ish, each correct with probability p, the vote is correct whenever the binomial tail exceeds half:

A_vote(n) = P(Bin(n, p) > n/2) = Σₖ₌⌊n/2⌋₊₁ⁿ C(n,k) · pᵏ · (1−p)ⁿ⁻ᵏ error ≈ exp(−n · D(½ ‖ p)) ← Chernoff: error decays exponentially in n

Majority voting wins when errors are diverse (different failure modes per sample — reasoning paths disagree) and p > ½-ish; it fails when errors are correlated (the model systematically misreads the problem — every sample makes the same mistake, and the vote ratifies it). Diversity across samples is why doc 14's temperature matters here: temperature 0 gives you n identical votes.

# best-of-n skeleton — the whole technique in 12 lines def best_of_n(problem, n=8): cands = [sample(problem, temperature=0.8) for _ in range(n)] # parallel scored = [(verify(c), c) for c in cands] # checker, not vibes return max(scored)[1] # cost: n × full_decode · accuracy: 1−(1−p·q)^n · fails iff p·q ≤ p

06 Reasoning-Model KV & Context Mechanics

A 30K-token thinking chain is 30K KV cache entries you pay for on every subsequent decode step.

Doc 06 established the anatomy: every generated token attends over the full KV cache, so decode cost per step grows linearly with context length — and doc 08 showed the binding constraint is memory bandwidth, not FLOPs. Reasoning models push all three levers at once:

Memory pressure. KV bytes = 2 × layers × kv_heads × head_dim × dtype_bytes × tokens. At ~0.1–0.5 MB per 1K tokens per user (model-dependent), a 30K-token reasoning chain is 3–15 MB of cache held for the whole request — thousands of concurrent users multiply this into the gigabytes that cap batch size.
Bandwidth tax per step. Every thinking token reads the entire accumulated cache. Token 30,000 of a chain decodes measurably slower than token 100 — long reasoning is O(L²) total bandwidth for L tokens. This is the wall from doc 08, met from a new direction.
Compaction returns. The fix is familiar from doc 07's economics and previewed in doc 19: summarize or prune the chain of thought mid-generation ("context compaction"), keeping a distilled state instead of the full transcript. Reasoning models trained with compaction tolerate it; naive ones degrade — the chain is the model's working memory, and truncating it amputates the plan.
🧠
The reframe: a chain of thought is RAM you rent by the token. Attention over past thinking tokens is how the model holds intermediate results — which is exactly why "the model forgot its earlier constraint" and "the response got slow and expensive" are the same bug with two symptoms.

07 When It Helps, When It Doesn't — and the Economics

Why RLVR beats naive RLHF for reasoning. Training long chains of thought needs a reward that can't be gamed. Verifiable reward r = 1[answer correct] — unit tests, exact answers, formal proofs — ties reward to reality. A learned human-preference reward model instead scores plausibility, and a long confident-sounding chain optimizes plausibility: reward hacking. RLVR's reward is cheap, exact on its domain, and unhackable within it — at the cost of working only where checkers exist. That boundary is precisely the table below.

DomainVerifier exists?Test-time compute payoff
Math / proofsExact (checker, formal system)Large — AIME-style jumps of 2–4× from thinking budgets
CodeExact (tests, compilation)Large — self-generated tests let the model verify itself
Science / analysisPartial (units, consistency checks)Moderate — helps structured reasoning, can't check facts
Summarization / NER / chatNo — reward is fuzzy human preferenceNear zero or negative — empirically overthinking: accuracy flat, cost 10–50×, sometimes worse

The economics. Thinking tokens are billed as output tokens — the 3–5× row of doc 07's table. The number that actually governs your purchase decision is cost per correct answer:

cost_per_correct = (avg_thinking_tokens × price_out) / p(correct) reasoning wins ⟺ cost_per_correct(reasoning) < cost_per_correct(baseline)

Worked example: baseline gives 60% correct at 400 output tokens; a reasoning model gives 92% at 6,000 tokens. Baseline: 400/0.60 ≈ 667 tokens per correct answer. Reasoning: 6,000/0.92 ≈ 6,522. If a correct answer is worth more than ~10× a wrong one (an escalation, a human retry), reasoning wins; if wrong answers are cheap to catch and retry, it doesn't. Extra thinking is not worth it when: p is already high (nothing to gain), the task has no verifier signal to steer the chain, errors are systematic rather than stochastic (majority-vote logic fails), or latency is the product.

✓ Do

Spend thinking tokens on verifiable, multi-step problems; expose "reasoning effort" as a per-task config; measure pass@1 and cost-per-correct, not vibes; batch parallel samples when the checker is cheap.

✗ Don't

Route extraction, classification, or summarization through a reasoner; pay for 30K thinking tokens on a retry-able 100-token task; assume longer thinking = better (the α < 1 curve flattens).

08 Mental Models

Scratch paper during an exam

The student (frozen brain — frozen weights) gets extra scratch paper and more clock time instead of a smarter brain. Hard problems improve a lot; "what is 7×8" doesn't. Lets you reason about: why the payoff concentrates in multi-step verifiable domains, and why easy tasks just get slower and pricier.

Scratch paper can't add knowledge the student never had — test-time compute refines search, it doesn't import facts (hallucination needs doc-level grounding, not thinking tokens).
Renting a bigger workshop vs. building one

Parameters are a factory you must build (capex, months, gating by capital); test-time tokens are a workshop you rent by the hour (opex, instant, per-project). Lets you reason about: the asymmetry of section 03, and why budget flexibility — 10× spend on the one hard problem — often beats a uniformly bigger factory.

Renting scales linearly in price with usage; a factory amortizes. At extreme, constant volume, owned capacity wins — which is why labs still build big models too.

09 Common Misconceptions

"Reasoning models are a different architecture." No — same transformer, same decode loop (doc 06). What changed is post-training: RL on verifiable rewards taught the model to use its own tokens as search. "Reasoning" is a behavior, not a module.

"More thinking tokens always mean better answers." The scaling curve J(c_test) ∝ c_test^α has α < 1 and flattens; beyond a problem-dependent point you pay 2× tokens for +1% accuracy. And on non-verifiable tasks, overthinking is empirically negative: more room to rationalize a wrong reading.

"Best-of-n / majority voting is obsolete now that models reason." It composes. Reasoning models raise per-sample p; sampling multiplies it — 1−(1−p)ⁿ still applies, just starting from a higher p. Parallel sampling is also latency-flat (batch), while serial thinking is not.

"Thinking tokens are free because they're hidden." They're billed as output tokens at full rate, they occupy KV cache for the request's lifetime (section 06), and they add latency linearly. Hidden from your eyes is not hidden from your invoice.

"Self-correction means I can trust the answer." Verification is only as good as the model's ability to check the domain. On verifiable tasks it's strong (tests, arithmetic); on factual recall, a model "checking itself" often re-confirms the same hallucination with more confidence.

🗺️
Where this goes next: long chains stress the KV cache exactly like long documents do — same wall, same fixes. Doc 19 covers the memory architectures (compaction, retrieval, state compression) that both problems converge on. And the economics here sit directly on doc 07's token-as-currency table: thinking tokens are the most expensive row, spent deliberately.
🎯
Closing: test-time compute is the first scaling axis where you hold the dial. The skill it adds to context engineering is allocation: match the token budget to the problem's verifiability and difficulty, measure in cost-per-correct, and remember that thinking is RAM you rent — spend it where a checker can steer it.