In a groundbreaking development that could reshape the economics of artificial intelligence, researchers have trained an AI model that achieves near-full performance while using only 12.5 percent of its experts. This milestone, reported by the-decoder.com on May 16, 2026, marks a significant leap in making AI more efficient, affordable, and accessible. For businesses, technologists, and society at large, this breakthrough isn't just about a single model—it signals a fundamental shift in how we think about AI's future: less is often more, and that changes everything.
Imagine a powerful AI that can answer questions, generate text, or help with complex tasks, but instead of needing a massive data center full of energy-guzzling hardware, it can run on a fraction of that power. That's the promise of this research. The model was trained to activate only a tiny percentage of its internal "experts"—the specialized parts of the network that handle different types of problems—yet it still performs almost as well as models that use all their resources. This isn't a small tweak; it's a rethinking of how AI models are built and deployed.
In this article, we'll unpack what this means for the future of AI, why efficiency matters more than ever, and how businesses and everyday users can benefit. Let's dive into the details.
Modern AI models, like the ones powering chatbots, image generators, and virtual assistants, have grown enormous. Some have hundreds of billions of parameters—the building blocks that store knowledge. While these models are incredibly capable, they come with a steep price tag. Training them requires vast amounts of electricity and specialized hardware, often costing millions of dollars. Even running them after training, known as inference, burns through energy and computing power with every query. For many smaller companies or developing nations, this cost barrier keeps cutting-edge AI out of reach.
The traditional fix has been to build bigger models to get better results. But the researchers behind this new work took a different approach. Instead of making the model physically smaller, they made it smarter about which parts to use. The key insight: most of a model's "experts" sit idle during any given task. Why wake them all up if only a few are needed? By training the model to rely on just 12.5 percent of its experts, they slashed resource use without sacrificing performance. This is a huge win for efficiency.
The report from the-decoder.com describes a model based on a technique called mixture-of-experts (MoE). In MoE models, different parts of the network specialize in different types of knowledge—like one expert for science, another for language, and another for reasoning. When the model receives a question, it only activates a few of these experts. The challenge has always been balancing which experts to use and how to train them to cooperate. This new research cracked that code.
The team trained the model to hit near-full performance—meaning scores comparable to larger, full-expert models—while using only 12.5 percent of its experts. This was achieved through a novel training method that forces the model to become extremely efficient at selecting and combining the few experts it needs. Think of it like a team of specialists where only two or three are called in for any given problem, yet they collaborate so well they match a much larger team's output. The result is an AI that's not only cheaper to run but also faster because less computation is wasted.
This isn't just an academic exercise. It's a practical blueprint for building AI that can run on everyday devices like phones or laptops, or in data centers using far less energy. The implications are enormous.
This breakthrough points toward a future where AI becomes dramatically more accessible. Here's what it means for different areas of technology and society.
One of the biggest barriers to AI adoption is cost. Large models require expensive cloud infrastructure, which prices out small businesses, nonprofits, and educators. By showing that a model can achieve near-full performance with just 12.5 percent of its experts, the researchers open the door to AI that runs on cheaper, less powerful hardware. Imagine a school in a rural area using an AI tutor that operates on a single server instead of a supercomputer. Or a local clinic using an AI diagnostic tool without needing a massive IT budget. Efficiency brings AI to the masses.
Data centers that train and run large AI models consume staggering amounts of electricity, often from fossil fuels. A model that uses only 12.5 percent of its experts could cut energy use by nearly 90 percent for each query. This is a game-changer for sustainability. As companies face growing pressure to reduce their carbon footprint, efficient AI models offer a way to maintain capability while slashing environmental harm. The future of AI isn't just about being smarter; it's about being greener.
Another exciting possibility is the rise of on-device AI. Smartphones, smart home devices, and even cars could run powerful AI models locally without needing the cloud. That means faster responses, better privacy (since data stays on the device), and offline functionality. A model that needs only a fraction of its experts can fit into limited memory and processor space. This breakthrough brings us closer to truly intelligent devices that work instantly and privately.
For business leaders, this research isn't just a tech headline—it's an opportunity. Here are actionable insights you can apply today.
Of course, no breakthrough is perfect. While the model achieves near-full performance, that "near" still means there's a slight trade-off. In some edge cases, the full model might still outperform it. The researchers are clear that more work is needed to close that gap entirely. Additionally, training this kind of efficient model requires careful design and specialized techniques that aren't yet standard. Most companies won't build their own models from scratch but will rely on providers to integrate this approach.
There's also the question of scale. This breakthrough was demonstrated with a specific model. Whether the technique works as well for even larger or different kinds of models remains to be seen. But early results are promising, and the trend toward efficiency is clear.
This story from the-decoder.com is part of a larger movement in AI. For years, the mantra was "bigger is better." But as models hit practical limits—both in cost and performance—researchers are turning to smarter architectures. Efficiency isn't a downgrade; it's a new frontier. By using only 12.5 percent of its experts, this model doesn't just save resources; it proves that AI can be both powerful and practical.
For society, this means AI will likely become more embedded in everyday life. From smarter traffic lights that reduce congestion to personal assistants that run entirely on your phone, the benefits will touch everything. But it also means we need to think about equity. Will efficient AI be made available to everyone, or will it remain behind corporate paywalls? The potential is there for broad positive impact, but it requires intentional action.
The research highlighted on May 16, 2026, by the-decoder.com is more than a technical achievement—it's a vision of a future where AI is lean, accessible, and sustainable. Hitting near-full performance with just 12.5 percent of an AI's experts shows that we don't need to keep building bigger to get better. For businesses, this means lower costs, faster responses, and new possibilities. For society, it means AI can reach more people while respecting our planet's limits. As this approach matures, expect to see a wave of smarter, greener, and more affordable AI tools. The age of efficiency has arrived, and it's a game-changer.