Um momento
0xB0Lesson 12 of 16

Reasoning models and test-time compute

Trade tokens for brains: why models that “think” first get harder problems right, and when the extra thinking is a waste of money.

25 min 7-question quiz 2 code exercises
By the end of this lesson you can
  • Explain why spending more tokens at inference time can improve answers
  • Compare sequential thinking with parallel sampling plus voting or a verifier
  • Weigh the cost, latency and faithfulness of reasoning traces

A Transformer does a fixed amount of computation per token. Ask for the answer to a tricky problem as the very first token, and the model gets one forward pass to “think”. Let it write out intermediate steps, and every step is another forward pass that can build on the previous ones. Chain-of-thought works partly because more tokens = more computation.

Pretraining made models bigger to make them smarter - train-time compute. Reasoning models add a second dial: test-time compute, spending more effort on each question when it’s asked.

How reasoning models are made

Models like OpenAI’s o-series and DeepSeek-R1 are trained with reinforcement learning on verifiable rewards: give the model thousands of math and coding problems whose answers can be checked automatically, let it write long reasoning before answering, and reward only correct final answers. Nobody hand-writes the reasoning style. Over training, the thinking traces grow longer and start to show behaviours like double-checking, backtracking (“wait, that’s wrong…”) and trying another approach.

Many APIs now let you set a reasoning effort or thinking budget: low, medium, high, or a maximum number of thinking tokens. The thinking tokens are usually billed as output tokens even when you can’t see them.

thinking_cost.py
1price_per_million = 10.0     # USD for output tokens
2speed = 80                   # tokens per second
3for effort, tokens in [("low", 300), ("medium", 2000), ("high", 12000)]:
4    cost = tokens / 1e6 * price_per_million
5    print(f"{effort}: {tokens} tokens, {cost:.3f} USD, {tokens / speed:.1f} s")
Output
low: 300 tokens, 0.003 USD, 3.8 s
medium: 2000 tokens, 0.020 USD, 25.0 s
high: 12000 tokens, 0.120 USD, 150.0 s

Thinking longer vs. thinking wider

There are two ways to spend test-time compute:

  • Sequential: one long chain of thought that revises itself. Great when steps depend on each other.
  • Parallel: sample N independent answers, then pick one. Self-consistency takes a majority vote; best-of-N asks a verifier (unit tests, a math checker, or a trained reward model) to score each answer and keeps the best.

If one sample is right with probability pp, the chance that at least one of N samples is right is 1−(1−p)N1 - (1-p)^N - this is why “pass@k” coding scores climb so fast. But you only get that answer if your verifier can recognize it. Majority voting needs no verifier but fails when the model is consistently wrong.

best_of_n.py
p = 0.3
for n in [1, 2, 4, 8]:
    print(f"N={n}: P(at least one correct) = {1 - (1 - p) ** n:.2f}")
Output
N=1: P(at least one correct) = 0.30
N=2: P(at least one correct) = 0.51
N=4: P(at least one correct) = 0.76
N=8: P(at least one correct) = 0.94

Try it

Worth thinking about?

Reasoning effort costs time and money. Sort each task: does it deserve extra thinking, or is a fast answer fine?

0 of 8 sortedScore 0/0
  • “A multi-step math word problem”

  • “Finding a bug that spans three files”

  • “Planning a database schema migration with zero downtime”

  • “A logic puzzle about who sits next to whom”

  • “Translating “thank you” into Spanish”

  • “Autocompleting the next few words in an editor”

  • “Labeling an email as spam or not spam”

  • “Naming the capital of Japan”

Key takeaways

  • More generated tokens means more computation - reasoning trades tokens, time and money for accuracy.

  • Reasoning models are trained with RL on checkable problems; effort settings control how long they think.

  • Parallel sampling helps when you can vote or verify; 1 − (1 − p)^N grows fast, but only a good verifier picks the winner.

  • Thinking traces can be unfaithful, and easy tasks don’t need them.

Lesson quiz

7 questions · pass with 5 correct · up to 50 XP

Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.

Practice: write Python

Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.

Exercise 1

Verifier vs. majority vote

+25 XP

Line 1 is the correct answer. Each following line is one sampled answer and its verifier score: answer score.

For every N from 1 to the number of samples, pick the highest-scoring answer among the first N (ties go to the earlier sample) and print N=1: 41 (0.62) wrong (score to 2 decimals, then right or wrong).

Finally do a majority vote over all samples - ties go to the answer that appeared first - and print majority: 41 (3/5) wrong.

  • The verifier beats the crowd
  • A fooled verifier
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Exercise 2

Pick a thinking budget

+25 XP

Line 1 is price_per_million tokens_per_second target_accuracy. Each following line is one effort setting: effort thinking_tokens answer_tokens accuracy.

All tokens (thinking + answer) are billed at the price and generated at the speed. For each setting print low: 20.0 s, 0.0100 USD, accuracy 74%. Then print pick: EFFORT - the cheapest setting that reaches the target accuracy - or pick: none - raise the budget or change models if none does.

  • Medium is the sweet spot
  • Nothing is good enough
main.py
Loading editor…

Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.

Questions about this lesson

Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.

Loading posts…

Gostou da aula? 😆👍
Apoie nosso trabalho com uma doação: