Blog
3 min read

What Is a Reasoning Model? AI That Thinks Before It Answers

Reasoning models spend extra computation working through a problem step by step before answering, which makes them better at maths, coding and multi-step tasks — at the cost of time and tokens. How they work, extended thinking and effort levels, and when to use one.

A reasoning model is an AI model that works through a problem — writing out intermediate steps, checking itself, trying alternatives — before giving its final answer. You'll also see this called "thinking", "extended thinking" or "chain of thought".

The basic idea

A standard language model produces its answer one word (token) at a time, straight away. For a simple question, that's fine. For "find the bug in this 300-line function" or a tricky maths problem, answering immediately is like doing long division in your head without paper.

A reasoning model gets the paper. It first generates a stream of reasoning — breaking the problem down, testing ideas, noticing mistakes — and only then writes the answer. More thinking usually means better answers on hard problems. (What is an LLM?)

Where this came from

People noticed early that simply adding "let's think step by step" to a prompt improved accuracy on maths and logic. This chain-of-thought prompting works because the model's own written-out steps become context for what it writes next.

Reasoning models build this in: they're trained to reason well, and many let you control how much thinking they do.

How you control it

Most current models expose some version of:

  • Thinking on/off or a thinking budget — how many tokens it may spend reasoning.
  • Effort levels — low, medium, high and so on. Higher effort means more thinking, slower and more expensive answers, and usually better results on hard tasks.

In Claude Code, for example, you can change the effort level with /effort. (How to change the model in Claude Code)

The trade-offs

Better on:

  • Debugging and complex code changes
  • Maths, logic and planning
  • Tasks with many steps or constraints
  • Analysing long documents for something specific

Costs:

  • Slower — it may think for many seconds before replying.
  • More tokens — thinking tokens are billed as output, which is the expensive kind. (What are tokens?)
  • Overkill for simple jobs — classifying a support ticket doesn't need it.

When to use it

Task Reasoning?
Summarise an email No
Classify or extract fields Usually no
Write a simple function Low effort
Fix a bug across several files Yes
Plan an architecture change Yes, higher effort
Solve a logic or maths problem Yes

A useful rule: start with low or default effort, and raise it only when the answers aren't good enough. For apps, test both on your real inputs and compare quality against cost. (Evaluating LLM outputs)

Can you see the thinking?

Sometimes. Some products show the reasoning, a summary of it, or nothing at all. Either way, don't treat visible reasoning as a guarantee — a model can reason its way confidently to a wrong answer. Check results that matter. (Why does AI hallucinate?)

The summary

  • Reasoning models think step by step before answering.
  • Much better on hard, multi-step problems; slower and pricier.
  • Control it with thinking budgets or effort levels.
  • Use it where it pays off; skip it for simple tasks.

EasySpawn runs Claude Code on a persistent server, so long, high-effort tasks can keep thinking and working after you close your laptop. See how it works or join the waitlist.

Related: What Is an LLM? · Claude Opus vs Sonnet vs Haiku · What Is Prompt Engineering? · LLM Temperature Explained

Keep reading