Anthropic engineer explains why Claude's writing got worse although the model got smarter

Why Claude's Writing Got Worse Even as the Model Got Smarter

By · Published September 23, 2026 · Updated September 23, 2026

Something strange has been happening inside the world of AI, and an Anthropic engineer has finally put a name to it. Claude, the AI assistant built by Anthropic, has been getting measurably smarter. It solves harder problems. It handles longer documents. It reasons through tougher questions. And yet, a large number of users have been quietly complaining about the same thing: the writing feels worse.

That sounds like a contradiction. How can a model get better at thinking and worse at writing at the same time? The answer turns out to be one of the most important lessons in modern AI, and it has huge implications for every business that is betting on these tools.

The Odd Truth: Smarter Doesn't Mean Better

When most people hear that an AI model is "smarter," they imagine a straight line. Better scores on tests, better answers, better writing, better everything. But AI models are not students who simply learn more. They are systems trained to hit specific targets. Change the targets, and you change the behaviour, sometimes in ways nobody intended.

That is the heart of the explanation coming out of Anthropic. The gap between what the model scores on tests and what users actually feel when they read its output is not a bug in one update. It is a structural problem in how AI is built and measured. The model got better at the things the company was measuring. The writing suffered because writing quality is not what those measurements were really tracking.

This is a much bigger deal than one chatbot sounding a bit stiff. It tells us something fundamental about the road ahead for the entire AI industry.

The Yardstick Problem: You Get What You Measure

There is an old business saying: what gets measured gets managed. In AI, it is more brutal than that. What gets measured gets optimized, whether you meant to or not.

AI labs train models using scores. Those scores come from benchmark tests, human raters, automated graders, and safety reviews. Press a model to win at those tests and it will find a way. Sometimes that way looks like genuine improvement. Sometimes it looks like learning the shape of the test instead of the substance of the skill.

This is often called Goodhart's Law: when a measure becomes a target, it stops being a good measure. A model can climb the leaderboard while drifting away from what real people find useful, clear, or pleasant to read.

Why writing quality is so easy to lose

Writing is one of the hardest things to score automatically. A grader can check whether an answer is factually right. It has a much harder time judging whether a paragraph flows, whether the tone fits, whether the rhythm feels human, or whether the whole thing is three times longer than it needed to be.

So when training rewards "correct" and "safe" and "complete," a model can slide toward prose that is technically fine but flat. Think of it like a restaurant being graded only on food safety. The kitchen passes every inspection, and the food slowly stops tasting like anything.

Three Forces That Pull Writing Downhill

1. Feedback from people is noisy

AI models learn a lot from human preferences. Raters pick which of two answers they like better, and the model adjusts. The trouble is that human preferences are messy. Raters are often rushed. They tend to favour longer answers because long feels thorough. They favour confident answers because confident feels smart. They favour answers that agree with the questioner because agreement feels helpful.

Train on millions of those choices, and you can accidentally teach a model to be wordy, over-confident, and eager to please, three habits that make writing worse while making scores go up.

2. Safety and helpfulness tuning shapes the voice

Before a model reaches the public, it goes through rounds of tuning to make it safe, polite, and cautious. This work matters enormously. But caution has a sound. It sounds like hedging. It sounds like "it depends." It sounds like a dozen disclaimers wrapped around a simple answer.

A model that is careful about every possible edge case will often write like a lawyer, not a writer. Safety tuning raises the floor on behaviour. It can also lower the ceiling on style.

3. Benchmarks reward the wrong kind of smart

Most public AI benchmarks test reasoning, math, coding, and factual recall. Very few test voice, clarity, rhythm, or usefulness in a real document. So a lab can post a big gain on reasoning and a real loss in writing, and its dashboards will look like nothing but green arrows.

The engineer's explanation lands on exactly this tension: a model can be genuinely more capable and still be less pleasant to work with. Both things are true at once, and only one of them shows up on a chart.

Why This Matters More Than It Sounds

For years, the AI story has been told as a simple upward curve. Each new model is bigger, faster, smarter. Users are supposed to feel the difference immediately. This episode complicates that story in a useful way.

It shows that capability is not one thing. It is a bundle of skills that can move in different directions. A model can get better at solving a tricky logic puzzle while getting worse at writing a warm email to a customer. If your business only measures the puzzle, you will be blindsided by the email.

It also shows that the people closest to these systems are willing to say so out loud. That kind of honesty is a sign of a maturing industry. The first phase of AI was about proving what was possible. The next phase is about proving what is actually good.

What This Means for Businesses

If you run a company that uses AI for writing, marketing copy, support replies, reports, proposals, here is the uncomfortable part. Your AI vendor's benchmark scores tell you almost nothing about whether the output will sound right to your customers.

That means you have to build your own tests. Not abstract tests. Real ones, using your real work.

What This Means for Society

There is a bigger picture here too. Most of us will not read benchmark tables. We will judge AI by how it feels to talk to. If the tools get technically stronger while feeling colder, more verbose, and more evasive, public trust will not follow the charts upward.

That is a real risk. AI is being woven into education, healthcare, customer service, and government. In all of those places, how something is said matters as much as what is said. A perfectly accurate answer delivered in a flat, robotic, over-hedged way can still fail the person reading it.

The good news is that this is a solvable problem, but only if the industry treats writing quality as a first-class goal, not a side effect. That means new kinds of evaluations, better human feedback, and rewards that value brevity and warmth instead of just length and confidence.

Actionable Takeaways

What to Watch Next

The real test is whether labs start shipping models that are measured on how they write, not just how they score. If evaluation improves, model behaviour will follow, because models become whatever their targets reward.

Watch for more honest reporting about trade-offs. Watch for evaluations that grade clarity and tone. And watch for whether the next generation of models feels like an upgrade in daily use, not just on a chart.

The lesson from this moment is simple and a little humbling. In AI, getting smarter is not the same as getting better. And the only way to know which one you have is to look at the work the model actually does, not the number at the top of the leaderboard.

TLDR: An Anthropic engineer has explained why Claude's writing can get worse even as the model gets smarter. The reason is that AI models become whatever their training targets reward, and today's targets favor test scores, safety, and length over tone, clarity, and voice. For businesses, this means vendor benchmarks say little about real writing quality, and every team using AI for content needs its own quality tests, version tracking, and human review. For the wider AI race, it is a signal that the next phase of progress will be judged not by charts but by whether the tools actually feel better to use.