On your internal math benchmark, accuracy versus average thinking token budget looks like: 1k tokens gives 42%, 4k gives 58%, 16k gives 66%. Each 4x budget increase is buying roughly half the previous gain. If the same diminishing trend continues, roughly what accuracy do you expect at 64k tokens?
Each 4x budget step halves the gain (+16, +8, then +4), so 16k at 66% extrapolates to roughly 70% at 64k thinking tokens.
Imagine learning a new language. The first month you go from zero to holding a basic conversation, huge gain. The second month you go from basic to comfortable, smaller gain. The third month you go from comfortable to confident, smaller still. Each new chunk of effort buys less than the chunk before it. Test-time compute on a reasoning model has the same shape. The first thousand thinking tokens move accuracy a lot. The next four thousand move it less. The next sixteen thousand move it less again. So if your gains are halving with each step, you can extrapolate the next step and see roughly where things land.
Concept explanation~2 min read
Everything you need to truly understand this topic: intuition, mechanics, step by step explanation, code, formulas, and worked example. Click to expand.
Concept explanation~2 min read
Everything you need to truly understand this topic: intuition, mechanics, step by step explanation, code, formulas, and worked example. Click to expand.
Test time compute scaling is the central economic story of reasoning models. The promise is that you can buy accuracy by spending more thinking tokens; the reality is that the price of each new accuracy point keeps climbing. A handful of data points is usually enough to read the shape of the curve, and that shape tells you when to stop paying for more thinking and switch strategy.
The stem hands you three data points: 1k -> 42%, 4k -> 58%, 16k -> 66%. The arithmetic on the per-step deltas yields a clean answer for 64k. The interesting part is what the shape implies about production decisions.
Reading the per-step deltas
Absolute accuracy is the wrong number to extrapolate from. The shape lives in the deltas.
From 1k to 4k tokens, accuracy moves from 42% to 58%. That is a gain of +16 points for a 4x budget increase.
From 4k to 16k tokens, accuracy moves from 58% to 66%. That is a gain of +8 points for the next 4x. The gain halved.
From 16k to 64k tokens, if the same halving pattern continues, the gain should be roughly +4 points. Add to 66% and you land near 70%.
This is a textbook saturating curve. In log-budget space (where each step is a 4x increase) the gains decay geometrically: 16, 8, 4. The continuous version of this is accuracy(budget) ~ a - b * budget^(-k) for some decay exponent k; the discrete halving per step pattern fits a curve where each multiplicative step in budget yields a multiplicatively smaller gain. The exact functional form does not matter for an extrapolation one step ahead; the pattern does.
Situations where this technique stops working.
2–4 min · Everything important, quickly.
Real products, models, and research that use this idea.
- OpenAI's 'Scaling test-time compute' research in 2024-2025 documented log-linear scaling curves with saturation on AIME and competition math.
- Claude Opus 4.7 thinking exhibits diminishing returns past a problem-dependent budget on internal benchmarks; effort levels are calibrated to the bend.
What an interviewer would ask next. Try answering before peeking at the approach.
QHow would you decide whether to spend another 4x on thinking versus switch to verifier-guided best-of-n?
Compare projected marginal accuracy per dollar between the two strategies; best-of-8 at moderate budget often dominates one chain at matched total compute when an oracle is available.
Red flags & common mistakes
The phrases that signal junior thinking. Click to expand.
Red flags & common mistakes
The phrases that signal junior thinking. Click to expand.
Assuming the early steep gains continue. The curve flattens fast; extrapolating linearly overstates how much accuracy more compute buys you.
60 second bullets to scan on the way to the call.
Compute per-step gains from a sequence of accuracy versus budget data points
Recognize a log linear with decay shape from the gains
Primary sources. Browse if you want the original framing.
Same topic, related formats. Practice these next.