Microsoft researcher builds a working neural network out of goats in Age of Empires II to critique AI science

The Goat-Powered Neural Network: How a Microsoft Researcher Used Age of Empires II to Expose AI’s Flaws

In June 2026, a story broke that made even hardened AI researchers chuckle. According to a report on the-decoder.com, a Microsoft researcher built a working neural network out of goats in the classic video game Age of Empires II. The goal wasn’t to create a new breakthrough in machine learning. Instead, the researcher wanted to critique the state of AI science itself – to show how easy it is to fool the metrics we use to measure progress.

At first glance, training a neural network using pixelated goats sounds like a joke. And in some ways, it is. But the stunt carries a deep message. It forces us to ask: Are we really making progress in AI, or are we just getting better at passing tests that don’t mean much? This article unpacks what the goat network means for the future of AI, how it will affect businesses, and what everyday people should take away from this unusual experiment.

What Exactly Did the Researcher Do?

Without giving away the full technical secrets (the source article doesn’t include a detailed breakdown), here is the general idea. Age of Empires II lets players control large herds of animals, including goats. By moving these goats around the map and using the game’s built-in pathfinding and actions, the researcher was able to simulate the behavior of artificial neurons. Just like in a real neural network, each goat could be thought of as a “node” that connects to others. The “weights” between nodes came from the distance the goats travelled or the way they responded to commands. The whole system could then learn simple patterns — essentially creating a working, if extremely slow, neural network inside a 1999 video game.

This is not the first time someone has used a video game to do real computation. People have built computers inside Minecraft and programmed calculators in Super Mario Maker. But this project stands out because it was done by a Microsoft researcher—someone inside one of the world’s leading AI labs—with the explicit goal of critiquing AI science.

Why a Goat-Powered Network Exposes a Real Problem

The core problem the researcher wanted to highlight is what AI experts call “benchmark gaming.” Almost every new AI model is tested against standard benchmarks—like ImageNet for image recognition or GLUE for language understanding. If a model scores higher than the previous best, the researchers claim they have made progress. But these benchmarks often have hidden flaws: they can be memorized, they don’t test true understanding, and they sometimes reward models that cheat in subtle ways.

By building a neural network out of goats, the researcher proved that you can achieve some level of “learning” in an absurd setup. The goat network might not beat GPT-5 on any real task, but it could pass some basic toy tests. And that’s the point: if such a silly system can claim to “learn,” then our measures of learning must be very weak. The researcher is essentially saying, “Look, the emperor has no clothes. Our benchmarks are not giving us the truth.”

This kind of critique is not new. For years, leaders like Yann LeCun and Jürgen Schmidhuber have warned that chasing numbers on artificial tasks leads to overfitting and fragility. But a visual, funny experiment gets the message across far better than a paper full of math. When you see actual goats moving around a medieval village to compute, you understand that the AI field might be relying on shaky foundations.

What This Means for the Future of AI Research

The goat network is a wake-up call. The future of AI research cannot depend solely on leaderboard positions. Instead, we need to develop richer evaluations that test robustness, generalization, and real-world usefulness. Here are three key shifts the experiment points to:

For businesses, the message is equally important. Companies that deploy AI for high-stakes tasks—like medical diagnosis, loan approval, or self-driving cars—cannot just buy a model with high benchmark scores. They need to understand its limitations. The goat network shows that even a model that passes tests might be fundamentally broken.

Practical Implications for Businesses and Society

For Tech Companies and AI Startups

If you are building an AI product, do not rely solely on academic benchmarks. They are often as silly as a goat-powered neural network. Instead, create custom tests that reflect your actual use cases. For example, if you are making a chatbot for customer support, test it with unusual, angry, or confusing customer queries, not just polite standard questions.

Also, invest in interpretability. Tools that show why your model made a decision can prevent catastrophic mistakes. If you cannot explain your model, do not trust it in critical applications.

For Investors and Executives

Demand evidence of real-world performance, not just leaderboard positions. When a startup says “our model achieved state-of-the-art on benchmark X,” ask what benchmark X measured. Was it a toy dataset from 2012? Could a goat network beat it? If so, the claim might be hollow. Look for demonstrations in messy, real environments.

For Regulators and Policy Makers

The goat network highlights the need for AI auditing standards. Just as cars must pass crash tests before hitting the road, AI systems should undergo independent evaluations that stress-test their safety and fairness. The European AI Act and similar frameworks are a start, but they need to be as rigorous as the tests the Microsoft researcher implicitly suggests: if a system cannot explain itself, or if it only works in a narrow video game world, it is not ready for society.

For the General Public

Be skeptical of AI hype. When you hear that “AI has achieved human-level performance,” remember that the goats can “learn” too. Many AI breakthroughs are impressive, but they are often brittle. Use AI tools for tasks like language translation or image generation, but verify any critical outputs. And support movements that demand transparency from AI companies.

Actionable Insights: What You Can Do Today

Whether you are a developer, a manager, or just a curious user, here are tangible steps inspired by the goat network:

Conclusion: The Goats Are Talking – Are We Listening?

Inside every AI system, there is a bit of the goat network. The same basic principles of weighted connections and activation functions drive both. But the goat network is a reminder that doing the math is not the same as understanding the world. A neural network made of goats can fit a curve, but it will never grasp what a goat actually is.

The future of AI depends on moving beyond numeric scores and towards meaningful evaluation. We need systems that can reason, adapt, and explain themselves. The Microsoft researcher’s playful experiment should not be dismissed as a joke. It is a mirror held up to the field, asking us to look at our own reflections. If we ignore the goats, we risk building AI that is impressive on paper but useless in practice.

So the next time you see a headline about an AI breakthrough, think about the goat network. Ask yourself: “Could a few dozen pixelated goats do the same thing?” If the answer is maybe, then maybe it is time to demand better science.

TLDR: A Microsoft researcher built a working neural network using goats in Age of Empires II to expose how weak many AI benchmarks really are. This experiment shows that current evaluation methods can be tricked by absurd setups, threatening real-world reliability. Businesses and society must push for more rigorous, real-world testing and interpretability in AI. The goat network is a powerful metaphor: if goats can learn, our measures of learning are not good enough.