When you pay for an AI model, you usually look at two numbers: the price per million tokens for input and output. These numbers feel solid, predictable. But a troubling trend is emerging across major AI providers – including Anthropic with its new Claude Sonnet 5 – where the headline token rate stays the same, yet the real cost goes up. This isn't a glitch; it's a deliberate pattern that changes how businesses should think about AI spending.
Anthropic, the company behind the Claude family of models, has long been seen as a reliable, ethical alternative to other AI giants. But with the release of Claude Sonnet 5, a concerning pattern continues: the company hides price increases behind unchanged token rates. This means the price per million input or output tokens stays the same as on the previous version, but the actual cost of getting the same quality of work or the same volume of output can rise significantly.
Let's unpack what's really happening, why it matters for the future of AI, and how you can protect your budget and your workflows.
At first glance, Claude Sonnet 5 charges exactly the same per-token rate as Claude Sonnet 4. That looks like a good deal – you're getting a "sonically" better version for the same money. But here's the catch: the model itself has been engineered in ways that increase the number of tokens needed to complete the same task. This can happen through several mechanisms.
First, the model might be prone to generating longer, more verbose responses. Even if it's smarter, it often adds extra clarifications, caveats, or unnecessary context. A simple answer to "Explain quantum computing in one sentence" might now take three sentences, consuming way more output tokens than before. For a business running thousands of queries a day, those extra tokens add up fast.
Second, the model's system prompt or internal instructions might have been changed to encourage more thorough analysis. While that sounds good, it can lead to token waste. If you ask for a code review on Claude Sonnet 4 and it produced 500 output tokens, the same request on Sonnet 5 might produce 800 tokens – with the same actionable result. You're paying 60% more in output costs for no extra value.
Third, tokenisation itself can change. AI models work on tokens, not words. A new version might use a different tokeniser that breaks longer words into more tokens. This is invisible to you but shows up on your bill. The input and output token counts on the API dashboard stay the same, but the actual text volume hasn't changed – the model just needs more tokens to represent the same text. That's a hidden price increase.
Anthropic has a pattern of doing this across major version bumps. It's not unique to them, but they have been particularly adept at it. The company may argue that the model is "more capable" and therefore worth the hidden extra cost. But for many businesses, the additional capability is unnecessary for their core tasks – they just want the same reliable output at a stable price.
The hidden-price-increase pattern represents a fundamental shift in the AI industry. After years of dramatic price cuts – remember when GPT-3 cost $0.06 per thousand tokens? – the market is maturing. Providers are under pressure to monetise their massive investments in compute and training. Instead of raising list prices (which would cause backlash), they're using technical changes to increase consumption per real-world use case.
This trend has three major implications for the future of AI.
First, the era of "cheaper and better" is over. For a while, every new model generation offered higher quality at lower or equal cost. That was driven by algorithmic efficiencies and economies of scale. But those gains are diminishing. Now, newer models often deliver only marginal capability improvements, and providers need to recoup costs. The hidden price increase is a way to keep the headline competitive while effectively charging more.
Second, AI procurement will need to become more sophisticated. Businesses can no longer rely on a simple per-token price comparison. They need to run their own benchmarks – measuring actual tokens consumed for specific, real-world tasks across different model versions. This is work, but it's the only way to spot hidden increases. In the future, we'll see third-party tools that audit model efficiency, showing not just cost per token but cost per task completed.
Third, we might see a pushback from users. As more developers and enterprises realise that the actual cost of work is rising, they may demand more transparency from providers. Some may start to customise their prompts to reduce verbosity – a practice that's already common, but now more critical. Others may explore self-hosted models or open-source alternatives that don't have the same incentive structure.
For the typical small to medium business using AI for customer support, content generation, or code assistance, these hidden increases can quietly eat into margins. Consider a company that uses Claude to answer product questions. They might notice that their monthly token spend hasn't changed, but the volume of queries they can handle for that spend has dropped because each answer consumes more tokens. Over time, they either pay more or serve fewer customers.
Enterprise customers with contracted prices may feel less immediate pain, but renewal negotiations will change. Providers might offer "stable token rates" but in practice the value per token degrades. Companies need to negotiate performance-based clauses – guaranteeing a maximum number of tokens per typical response, or linking price to actual output quality rather than just token count.
Society as a whole should be aware that AI costs are not as transparent as they seem. This lack of transparency could widen the gap between large and small players. Big tech companies can run extensive testing and negotiation; small startups might just accept the published token rate and get blindsided by cost creep. The industry needs clearer standardised benchmarks that measure cost per accurate answer, not just cost per token.
Furthermore, the environmental impact of hidden price increases shouldn't be overlooked. More tokens mean more compute, which means more energy. If providers are incentivising longer outputs without corresponding value, that's wasted computational resources. Sustainable AI requires efficiency – and hidden token inflation works against that.
So what can you do right now to protect your organisation from hidden AI price hikes? Here are five concrete steps.
1. Benchmark before you upgrade. When a new model version is released, don't blindly switch. Run a set of representative tasks – 20 to 50 queries that match your actual usage. Record the token counts and the quality of outputs. Compare those metrics with the previous version. If token count per task has increased by more than 5-10%, the model is effectively more expensive, even if the per-token price hasn't changed.
2. Use length constraints aggressively. Set maximum tokens in your API calls. Even better, use system prompts that explicitly ask for concise answers. For instance: "Respond in no more than 50 words" or "Provide only the code, no explanation unless asked." This can counteract the model's tendency to be verbose.
3. Monitor your effective cost per use case. Rather than just tracking total spend, track how many customer queries, invoices processed, or lines of code you generate per dollar. Create a dashboard that updates in real time. If the metric drops after a model upgrade, you'll spot the problem immediately.
4. Negotiate with providers. If you're a volume customer, demand pricing based on performance, not tokens. Ask for a guarantee that the average token count per standard task won't increase between versions without a corresponding price cut. Some providers may be willing to offer custom contracts with efficiency benchmarks.
5. Mix and match models. Don't use one model for everything. Use a cheaper, smaller model for simple tasks and keep the powerful but costly model for complex ones. That way, if the expensive model becomes bloated, you limit its exposure to only where you truly need its intelligence.
These strategies will help you maintain control over your AI costs even as providers continue the pattern of hiding price increases behind unchanged token rates.
The release of Claude Sonnet 5 is just one data point in a broader industry shift. We will likely see more of this from other major players. The days of always cheaper and better AI are behind us. The future will involve more complex pricing, with providers using model architecture, tokenisation, and response length adjustments as levers to influence effective costs.
This doesn't mean AI is becoming unaffordable. It means that users need to be more savvy. The businesses that thrive will be those that invest in measurement, optimisation, and negotiation. They will treat AI not as a magical utility but as a cost centre that requires careful management.
Ultimately, the pattern highlighted by Claude Sonnet 5 is a wake-up call. Providers are businesses too, and they will find ways to maximise revenue. Our job as consumers of AI is to see past the surface-level metrics and understand what we're really paying for. The tools and practices I've outlined here will help you do exactly that.
Stay curious, stay critical, and always measure what matters.