RAG & LLM Engineering · BraivIQ AI Engineering Playbook
The Price Of Thinking: Reasoning Effort, Thinking Budgets And How To Spend Reasoning Tokens In Production - A Senior Engineer's Guide
Reasoning models quietly broke the cost model every team had internalised. A standard completion returns an answer. A reasoning model first produces a thinking trace of four to sixteen thousand tokens - sometimes far more - and that trace counts against the context window, is billed as output, and routinely makes a request run an order of magnitude longer than a chat completion. Left on the default, the smartest models are also the most expensive and the slowest, and the gain is real only on the fraction of requests that actually needed the thinking. In 2026 the providers exposed the dial - OpenAI's reasoning effort levels, Anthropic's effort parameter that lets the model decide how much to think within a ceiling - and the practitioners who learned to use it report cutting cost by 40-60% with no loss on the work that matters. This educational deep-dive, for senior engineers and CTOs, explains what thinking tokens are and why they cost what they do, why effort is a model-specific contract rather than a universal setting, which tasks pay back the thinking and which burn it, and the production pattern that works: classify first, budget to match, default low, escalate on signal, and measure per task.
· 13 min read · By BraivIQ Engineering
4k-16k tokens - A typical reasoning trace per request - counted against context and billed as output · ~10x longer - Reasoning-mode requests routinely run an order of magnitude longer than standard completions · 40-60% - Cost reduction teams report from routing simple requests to low effort and reserving high effort for the hard ones · Classify, then budget - The 2026 production pattern: a cheap pre-classifier picks the effort level. Default low, escalate on signal
For most of the LLM era, cost was a simple function: tokens in, tokens out, a price per million of each. Reasoning models broke that model in a way many teams have still not internalised. Before a reasoning model answers, it thinks - it generates a trace of intermediate reasoning, typically four to sixteen thousand tokens and sometimes far more, that the user never sees. That trace counts against the context window, it is billed as output tokens, and it takes time to produce, which is why reasoning-mode requests routinely run an order of magnitude longer than a standard chat completion. The consequence is a quiet inversion: left on their defaults, the smartest models in your stack are also the most expensive and the slowest per request, and the gain from thinking is real only on the fraction of requests that actually needed it. A classification, a lookup, a formatting task gets the same expensive trace as a genuinely hard multi-step problem, and the difference shows up as a bill nobody can explain and latency nobody wanted. In 2026 the providers exposed the dial. OpenAI's reasoning effort takes levels from minimal through low and medium to high. Anthropic moved from a fixed thinking budget to an effort parameter - low, medium, high and beyond - within which the model dynamically decides how much thinking a given request needs. The practitioners who learned to use these controls report cutting cost by 40-60% while keeping accuracy on the work that matters. As an AI Agency Developer London that runs reasoning models in production for clients, we think the price of thinking is the most under-managed line in most AI budgets, and this educational deep-dive is how to manage it.
Effort Is A Model-Specific Contract, Not A Universal Dial
The first thing a senior engineer should internalise is that an effort setting is not a portable number. Research published this year makes the point precisely - reasoning effort is a model-specific API contract - and it matters in practice. OpenAI's levels control how much the model is permitted to reason, with minimal suppressing most of the trace for latency-sensitive work and high allowing extended deliberation. Anthropic's effort parameter sets a ceiling within which the model itself decides how much thinking each request warrants, so low effort on an easy request may think almost nothing while low effort on a hard one thinks more - the model is doing some of the classification for you. The same nominal level therefore produces different traces, different costs and different latencies across providers, and often across versions of the same provider's models. Treating effort as a global constant to be set once is the error that produces both over-spending and under-performing: the correct treatment is a per-task, per-model parameter, chosen by measurement, that lives in configuration alongside the model name and is re-validated whenever either changes. Put another way, the right effort level is a property of the pairing of a task and a model, not of your application, and your routing layer should hold it that way.
- Thinking pays - multi-step reasoning, planning, agent decisions with several options, hard extraction from messy or adversarial input, problems with traps.
- Thinking burns - classification, routing, lookup, formatting, summarising known content, anything a pattern match answers.
- Effort is per task and per model - a configuration value chosen by measurement, re-validated on every model change, never a global constant.
- Context is a cost too - a long trace consumes the window that retrieved documents and history needed. Budget for it.
- Latency is the hidden price - an order of magnitude slower per request changes what is viable in an interactive path.
The Production Pattern: Classify, Budget, Default Low, Escalate
The pattern that has emerged across teams running reasoning models at scale is simple to describe and pays for itself immediately. First, classify the request before you spend on it: a cheap, fast pre-classifier - a small model, a heuristic, a lookup against known request types - sorts incoming work into three or four buckets by expected difficulty, and most of the win comes from this single step, because it is what lets easy requests avoid the expensive path entirely. Second, map each bucket to a budget: the easy bucket goes to a non-reasoning model or minimal effort, the medium bucket to low or medium effort, and only the hard bucket to high effort - which is the hybrid routing that reports 40-60% cost reductions. Third, when the classifier is unsure, default low and escalate on signal rather than defaulting high: run the request at a low budget, check the result against your validation (does it parse, does it pass the schema, is the model's confidence or self-check acceptable), and only if it fails or signals low confidence re-run at a higher budget. Escalation on failure costs you a second call on the minority of requests that needed it. Defaulting high costs you the expensive path on every request that did not. Fourth, measure per task: build an evaluation set per request type and record accuracy, cost and latency at each effort level, so the budget mapping is evidence rather than intuition and so a model upgrade - which changes the contract - is re-measured before it ships. The research frontier is moving this inside the model, with certainty-guided and budget-aware reasoning that lets the model stop thinking when it is confident, but the operational pattern is available today, and it is the difference between a reasoning bill that scales with your hard problems and one that scales with your traffic.
The Bottom Line
Reasoning models changed the cost model: the thinking trace - four to sixteen thousand tokens and more, billed as output, occupying context, running an order of magnitude longer than a chat completion - is the most expensive thing most AI systems now do, and on defaults it is spent on every request whether or not the request needed it. The 2026 controls - OpenAI's effort levels, Anthropic's effort ceiling within which the model decides - are the instrument for spending it only where it pays, and the first discipline is to treat effort as a model-specific, per-task contract chosen by measurement, never a global constant. The production pattern is classify first with a cheap pre-classifier, budget each bucket to match, route easy work to non-reasoning or minimal effort and reserve high effort for the genuinely hard, default low and escalate on failed validation or low confidence rather than defaulting high, and measure accuracy, cost and latency per task so every mapping is evidence and every model change is re-validated. Done this way reasoning spend scales with your hard problems rather than your traffic, costs fall by 40-60%, and the hard cases get a better answer because the budget is concentrated where it matters. The price of thinking is real. Paying it only where thinking earns its keep is exactly the engineering we do.
References & Further Reading
- arXiv - the price of thinking: reasoning effort as a model-specific API contract: https://arxiv.org/pdf/2608.16956
- Boundev - reasoning effort: cut LLM cost and latency in production: https://www.boundev.ai/blog/reasoning-effort-llm-cost-latency
- Towards AI - LLM reasoning budget: how developers should spend thinking tokens without wasting latency: https://pub.towardsai.net/llm-reasoning-budget-how-developers-should-spend-thinking-tokens-without-wasting-latency-20ed2e43a31e
- arXiv - certainty-guided reasoning in large language models: a dynamic thinking budget approach: https://arxiv.org/pdf/2509.07820
- arXiv - BudgetThinker: empowering budget-aware LLM reasoning with control tokens: https://arxiv.org/pdf/2508.17196