Almost everything we trust in the digital world rests on a strange and wonderful trick: mathematics. When you type a password, send a message, or buy something online, your data is protected by encryption schemes that cryptographers have proven, on paper, with logic, cannot be broken without solving a problem so hard it would take more energy than exists in the universe.
Now imagine applying that same standard to artificial intelligence. Not "we tested it a lot and it seemed fine." Not "our red team tried to break it and couldn't." But an actual mathematical proof that a system will behave safely, the way a cryptographic proof guarantees a code cannot be cracked.
That is the ambition behind the Mathematical AI Safety Institute, which is pursuing a goal that could reshape how the entire AI industry earns trust: proving AI is safe the way cryptographers prove codes are unbreakable.
It sounds like an academic fantasy. But the consequences, for businesses buying AI, for regulators writing rules, and for anyone who will live alongside these systems, could be enormous.
To understand why this matters, you need to understand what a cryptographic proof is not. It is not a claim that a specific lock is unpickable. It is a claim that if a certain mathematical problem is hard, then breaking the lock is hard. The whole thing is conditional. It lives inside a carefully defined set of rules and assumptions.
Cryptographers call this a "reduction." You reduce the security of your system to a well-studied hard problem. If someone breaks your system, they've also solved the hard problem, and nobody believes that's possible. That logical chain is what lets the world run trillions of dollars of commerce through encryption that no human has ever physically verified by hand.
The AI safety field has nothing like this. Today, safety claims rest on:
All of these are valuable. None of them are proofs. They tell you what a model did in the past, on the inputs somebody thought to try. They say almost nothing about what the model will do tomorrow, on an input nobody has imagined yet.
This is the gap the Mathematical AI Safety Institute is trying to close, and it's a gap that has grown uncomfortable as AI systems move into medicine, finance, infrastructure, and military planning, where "it passed our tests" starts to sound alarmingly thin.
The appeal of the cryptographic approach is that it transfers trust out of the lab and into logic. You don't have to trust the vendor. You don't have to trust the tester. You don't even have to trust the institute. You can check the proof yourself.
That is a radically different basis for confidence than anything the AI industry currently offers. And it maps onto a real business problem: nobody knows how to verify AI safety claims independently. A company that wants to buy an AI system for, say, loan decisions or medical triage has to take the vendor's word for it. Auditors have no standard to audit against. Regulators have no yardstick to measure. Insurers have no way to price the risk.
A mathematical proof would change all of that. It would be portable, checkable, and, at least in principle, permanent. It would give a regulator something to require, a buyer something to demand, and a court something to examine.
Here is where the cryptography analogy gets genuinely difficult, and where anyone following this effort should pay close attention.
Cryptography works because the systems involved are small, fixed, and precisely specified. A block cipher does exactly one thing. Its behavior can be written down completely. There is no ambiguity about what it is supposed to do.
AI systems break all three of those properties:
Before you can prove a system is safe, you must define what "safe" means in precise mathematical language. For an encryption algorithm, the goal is crisp: nobody can recover the message without the key. For an AI system, the goals are things like "be helpful," "don't be harmful," "don't deceive," "respect user intent." These are human concepts. Converting them into formal statements without losing the meaning is one of the deepest unsolved problems in the field.
Prove the wrong property perfectly and you've achieved nothing, or worse, created false confidence.
Formal verification techniques have been applied successfully to software like compilers and operating system kernels, but those efforts took years and required teams of specialists for code far smaller than a modern AI model. Scaling mathematical proof to systems with billions of parameters is not a matter of trying harder. It likely requires entirely new techniques.
A cryptographic proof holds because the system operates in a closed world with defined inputs. An AI model deployed in the real world faces inputs nobody anticipated, in contexts nobody anticipated. Any proof must therefore say something meaningful about how the system handles the unexpected, which is precisely where intuition and testing fail most badly.
None of these obstacles make the goal impossible. But they do mean that any "proof of AI safety" will come with a long list of assumptions attached. Those assumptions are where the real debate will happen.
For most companies, the practical question is simple: will this change how I buy and deploy AI? Eventually, yes, and the shift will look familiar to anyone who lived through the arrival of security certifications.
Think about what happened with data security. It started as a matter of trust, became a checklist, then became a formal audit standard, and finally became a line item that legal teams demanded before signing a contract. AI safety is on a similar path, and formal methods are the most credible candidate for the "audit standard" stage.
What to watch for and prepare for:
The counterintuitive advice: don't wait for perfect proofs. Start now by documenting your safety claims precisely. The discipline of writing down exactly what you claim your system does, under what conditions, is the first step toward any formal argument, and it improves your engineering immediately.
The bigger prize is legitimacy. Public trust in AI is fragile, and it is fragile for a rational reason: the technology is powerful, opaque, and evaluated mostly by the people selling it. "Trust us, we tested it" is not a foundation a society can build on.
Mathematics offers something different. A proof doesn't care who wrote it. A skeptic in any country can check the logic. That universality is why cryptography became the backbone of global commerce despite deep geopolitical distrust between the parties using it. There is no reason to think international AI governance will be any different, verification that doesn't require trusting the other side is the only kind that scales across borders.
But society should also be clear-eyed. A proof is only as good as its assumptions, and those assumptions will be chosen by humans with interests. The failure mode here is not that AI safety proofs are impossible. It is that a weak proof gets marketed as a strong one, and the word "proven" does what the word "natural" did on food packaging.
That means the most important byproduct of this effort may not be the proofs themselves, but a culture of asking: proven under what assumptions? Guaranteed against what threat model? Checked by whom?
The AI industry is at the point where a technology crossing into every part of daily life needs a trust architecture, not just better products. We built one for the internet, encryption, certificates, protocols, standards bodies, and it took decades, crises, and painful lessons.
AI does not have decades. The systems are already in the wild, making consequential decisions, and the gap between what they can do and what we can prove about them is widening every year.
That is what makes the Mathematical AI Safety Institute's ambition so significant. It is a bet that the field can skip the painful path and go straight to the rigorous one, that AI can be trusted not because we tested it enough, but because we can prove what it will do.
Whether or not that bet pays off fully, it reframes the entire conversation. The question stops being "how do we feel about AI safety?" and becomes "what can we actually demonstrate?" That is a harder question, a more honest one, and the only kind that will hold up when the stakes get real.