Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits

Anthropic Says Open-Weight GLM-5.3 Nearly Matches Claude Mythos Preview at Building Exploits, What That Means for the Future of AI

By · Published September 30, 2026 · Updated September 30, 2026

For years, the AI world has run on a comfortable assumption. The most powerful models sit behind paywalls, APIs, and terms of service. Open-weight models, the ones anyone can download and run on their own machines, trail behind. Maybe by a year. Maybe by two. Everyone plans around that gap, and that planning has been mostly right.

A new claim from Anthropic puts a crack in that assumption. Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits.

That is one sentence, but almost every word in it matters. Open-weight. Nearly matches. Building exploits. Together they describe a shift that security teams, business leaders, and policymakers should all be paying attention to right now.

First, What Does "Building Exploits" Actually Mean?

An exploit is a piece of code or a set of steps that takes advantage of a weakness in software to do something the software was never supposed to allow. Think of a locked door. A vulnerability is the discovery that the lock is faulty. An exploit is the actual tool that opens the door.

That second step is the hard one. Finding bugs is hard enough. Turning a bug into a reliable, working attack is a different and often much harder skill. It takes deep knowledge of how systems behave when they break, and it takes a lot of trial and error.

For most of computing history, this skill lived in a very small number of people. Nation-state labs, elite researchers, a handful of specialists. It was expensive, slow, and gated by human expertise.

AI models that can do this work change the economics entirely. Instead of a scarce skill, exploit development becomes something closer to a service you can query. That is the shift Anthropic is flagging.

Why the Word "Open-Weight" Is Doing All the Heavy Lifting

There is a big difference between a closed model and an open-weight model, and it has nothing to do with how smart the model is.

A closed model lives on someone else's servers. You reach it through an API. If the model's maker decides you should not have access, you lose access. If you misuse it, they can watch, throttle, or ban you. The model is contained, at least in theory, by the company that runs it.

An open-weight model is different. The parameters, the learned numbers that make the model work, are published. Anyone can download them. Anyone can run them on their own hardware. Anyone can fine-tune them, adapt them, and strip away the safety training the original developers added.

Once weights are released, they cannot be recalled. They cannot be patched out of existence. They spread. Copies exist on laptops, servers, and cloud instances around the world, and no company can switch them off.

This is why the comparison matters so much. A closed frontier model that can build exploits is a serious concern, but it is a manageable concern. An open-weight model at nearly the same capability level is a different category of problem. The question stops being "how do we monitor this?" and becomes "what do we do now that it's out there?"

"Nearly Matches" Is the Part People Underestimate

It is easy to read "nearly matches" and relax. Nearly is not the same as matching. The frontier model is still ahead.

But in security, "nearly" is often more than enough. You do not need the single best tool in the world to break into a system. You need a tool that works. Attackers are not competing for a trophy. They are looking for one path in, and most real-world targets are not defended by the most advanced technology available.

There is also a volume effect. If a model can generate ten candidate approaches in the time it would take a human expert to sketch one, quantity starts to beat quality. Many attacks succeed because someone tried a thousand times, not because they had the perfect plan.

So the interesting number here is not the size of the remaining gap. It is the fact that the gap is small enough to describe as "nearly."

Why It Matters That Anthropic Is the One Saying This

Anthropic builds Claude. Claude Mythos Preview is its own model. So the company is essentially pointing at a competitor's freely downloadable model and saying it has almost caught up to its own.

That runs against normal marketing instinct. Companies rarely volunteer comparisons that make their products look less unique.

Read one way, this is a safety signal, a lab choosing to raise an uncomfortable technical finding rather than keep it internal. Read another way, it is an argument aimed at policymakers: if open-weight models are approaching frontier offensive capability, then the rules about who can release what, and under what conditions, may need rethinking.

Both readings can be true at once. What is not in dispute is that a major frontier lab has put the comparison on the record.

The Trend Behind the Headline

The bigger story here is not one model. It is the direction of travel.

Open-weight releases have been closing distance on frontier systems for a while now. Each generation of freely available models seems to land closer to the cutting edge than the last. The lag that everyone used to plan around keeps getting shorter.

That has enormous upside. Open models drive down costs, let companies keep data on their own infrastructure, and put powerful tools in the hands of researchers, startups, and smaller organizations that could never afford frontier API pricing at scale.

It also means that whatever capability advantage the frontier has is temporary by design. Any business strategy that assumes a permanent gap is building on sand.

What This Means for Security Teams

The practical takeaway for defenders is uncomfortable but simple: assume exploit development is getting cheaper and faster.

That single assumption reshapes priorities. If creating a working attack used to take weeks of expert time and now takes hours of model time, then the window between a vulnerability being discovered and being exploited shrinks. Patching cycles built for a slower world start to look dangerously long.

There is a flip side, and it matters. The same class of model that can write an exploit can also read one. Defenders can use AI to triage vulnerability reports, suggest fixes, scan code, and spot patterns humans would miss. The technology is not inherently on the attacker's side, it is on the side of whoever uses it well and quickly.

But the asymmetry is real. An attacker needs one working path. A defender needs to close every path. That imbalance has always favored offense, and cheaper exploit development makes it worse.

What This Means for Businesses

Most companies are not going to run their own exploit-generation evaluations. But they will feel the downstream effects.

Expect faster exploitation of known bugs across the software you depend on. Expect vendor patching to become a competitive differentiator. Expect cyber insurance and risk models to start pricing in the reality that attack tooling is abundant.

There is also an internal governance question that most organizations have not answered. Employees can now download powerful open-weight models onto their own laptops. Those models may have safety training removed. They may be running outside any company policy, any logging, any oversight. That is shadow AI in its most acute form, and it deserves a clear, written answer long before an incident forces one.

The Policy Problem Nobody Can Dodge

Governments have spent years trying to control advanced AI through levers like export controls and compute thresholds. Those tools work, to the extent they work, at the point of training, they try to stop a model from being built.

Open weights break that model of control. Once a model is published, the regulatory question shifts from containment to deterrence. You cannot recall the weights. You can only shape the decisions of the people who release them, through norms, licensing, liability, or agreed thresholds for what should not be released openly.

That is a much harder problem, and it is now front and center. If a freely downloadable model can nearly match a frontier system at building exploits, then "should this be released?" becomes a question with real-world consequences rather than a thought experiment.

Actionable Steps You Can Take Now

The Bigger Picture: Capability Is Spreading Faster Than Governance

Step back and the pattern is clear. Capability is diffusing faster than the rules, norms, and institutions meant to manage it.

Frontier labs push the boundary. Open-weight releases pull those capabilities into public hands, often within a much shorter window than anyone expected. Regulators write rules for a world where models stay behind APIs, then discover that world is already gone.

None of this means open-weight models are a mistake. They have driven down costs, expanded access, and put serious tools in the hands of people who would never otherwise have them. The benefits are real and large.

It means the tradeoffs are now concrete. A claim like the one Anthropic has made forces a question that the industry has been able to defer: at what capability level does an open release stop being a public good and start being a public hazard? There is no clean answer, and pretending otherwise does not help.

What is clear is that the era of comfortable assumptions is over. The gap between what the best labs can do and what anyone can download is narrowing. Businesses, security teams, and governments should plan for a world where that gap keeps shrinking, and where the tools once reserved for a handful of experts are available to everyone.

Conclusion

The claim that Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits is easy to read past. It sounds like a benchmark footnote. It is not.

It is a signal that offensive cyber capability is becoming downloadable, that the frontier's lead is temporary, and that the safeguards built for a world of gated APIs no longer fit the world we actually live in.

The organizations that come out ahead will be the ones that treat this as a planning input today, shortening their patching cycles, governing how open models are used inside their walls, and putting AI to work on defense as aggressively as attackers will put it to work on offense.

The gap is closing. The question is whether our defenses and our rules close it too.

TLDR: Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits, meaning a freely downloadable model has come close to frontier-level capability at creating real attacks. Because open weights can never be recalled once released, this shifts the problem from monitoring to coping. For businesses, it means exploit development is getting cheaper and faster, patching windows must shrink, and open-weight model use needs a clear internal policy. For governments, it means compute thresholds and export controls no longer contain the risk. The takeaway: the lead frontier labs hold is temporary, and the organizations that plan for a narrowing gap now will be the ones that stay secure.