AI models keep getting smarter on paper. Year after year, they answer harder questions, solve trickier problems, and write more polished prose than ever before. Yet anyone who actually spends time talking to them knows the numbers only tell part of the story. Sometimes a model feels genuinely brilliant. Other times it feels hollow, even when the scorecard says it should be impressive. That gap between "high score" and "great experience" is why a growing number of AI insiders are talking about something far less formal: the vibe test.
OpenAI co-founder Karpathy is at the heart of this conversation. A sharp, widely respected voice in AI discussions, he has been on the lookout for the next AI vibe test. And the clues surrounding the search are delightfully strange: a unicorn, a pelican, and the sprawling fantasy world of Middle-earth.
At first glance, that sounds like a fairy tale. But the hunt for the next vibe test is really a hunt for a better way to measure intelligence — one that takes imagination, personality, and plain human feeling seriously. Here is what that means for the future of AI and how we will all use it.
A vibe test is a quick, informal check of how an AI feels to talk to. There is no official score and no multiple-choice sheet. Instead, you just chat. You ask strange questions. You make odd requests. You watch how the model handles humor, tone, surprise, and rules that nobody wrote down. Then you trust your gut: Did it feel smart? Did it feel human? Or did it feel like a machine trying too hard?
Think of it like meeting someone for coffee. A resume tells you what they have accomplished. A vibe test tells you whether you would actually enjoy the conversation. Both matter, but they measure very different things.
In practice, a vibe test might sound silly. You could ask an AI to explain why a pelican would make a great travel agent. You could ask it to write a poem about a unicorn that wanders into Middle-earth. These prompts do not look like serious science, but they secretly test a great deal: creativity, logic, tone, cultural knowledge, and the ability to keep a story straight from beginning to end.
Those three words are not random decorations. Each one pressure-tests a different kind of intelligence, which is exactly why they make such good material for the next generation of AI evaluation.
Now combine the three. Ask an AI to imagine a unicorn and a pelican teaming up for a quest through Middle-earth, and you have forced the model to improvise with no rehearsed script. That is exactly where the spark of real intelligence shows up — or where its absence becomes obvious. Whimsy, it turns out, is a very serious test.
For years, the AI world tracked progress with standardized tests. Fill in the blank. Pick the right answer. Solve the equation. Those benchmarks were genuinely useful; they helped teams measure progress and compare one model against another.
But they have real limits. Modern models are trained on enormous oceans of data, and some of that data includes test-style content. Scores can creep upward even as everyday usefulness flatlines. Worse, a score tells you nothing about what a person feels in a two-minute conversation. A model can ace an exam and still feel cold, repetitive, or vaguely robotic when a user actually tries to talk to it.
That is the hole the vibe test fills. It is not meant to replace benchmarks — it is meant to remind the industry that the final judge of any AI is a human being with a feeling, not a grading sheet. As AI slides deeper into work, school, and home life, the experience of using it becomes just as important as the accuracy of its answers.
The search for the next vibe test is a signal that AI is entering a new era: the era of experience. We are moving from "Can it answer correctly?" to "What is it like to live and work with this thing?" That shift will reshape what companies build, what customers expect, and how AI is judged.
First, personality becomes a product feature. If charm, warmth, and wit can be measured — even informally — builders will put real effort into tuning them. The most successful AI assistants may not be the smartest ones on paper. They may simply be the ones people enjoy talking to.
Second, evaluation gets more human. Companies will blend hard numbers with structured "vibe testing" that captures delight, trust, and clarity. That means everyday users, not just engineers, will become the new quality inspectors. A grandmother's impression of an AI health assistant will matter as much as a developer's unit tests.
Third, creativity becomes part of the bottom line. If imagination is what makes a model impressive, creative uses of AI — storytelling, games, design, brainstorming — will grow fast. The same technology that helps a doctor review medical records may also write a bedtime story about a homesick pelican.
You might not be building the next frontier model. But if your business uses AI — a chatbot, a writing helper, a research tool — you need a way to tell good from great. Here is how to borrow vibe-test thinking in your own organization.
The whimsical trio of unicorn, pelican, and Middle-earth might look like a playful distraction from serious AI research. In truth, it is a healthy correction. Technology improves when people ask not only "Is this fast?" but "Is this genuinely good to be around?" AI is becoming a companion in our work, our classrooms, and our homes. The sooner we get good at measuring how it makes us feel — with a mix of hard numbers and good old-fashioned intuition — the sooner we build tools that truly serve us.
The next great vibe test will not just be a clever prompt. It will become a shared language for talking about the one thing benchmarks cannot capture: the spark. And we are learning to see that spark one conversation at a time — starting with a unicorn at a crossroads, a pelican with a travel bag, and the long, winding roads of Middle-earth.
Karpathy's hunt for the next AI vibe test is more than an inside-baseball curiosity. It is a sign that the AI world is growing up. Raw intelligence was the first chapter. Next comes something harder to measure and just as important: character. As models become part of everyday life — helping us write, code, learn, heal, and create — the question "How smart is it?" will slowly give way to "How does it feel to work with it?" The models that win the future will be the ones that pass both the math test and the vibe test.
The unicorn, the pelican, and Middle-earth are not just charming riddles. They are early clues about how we will judge machines in the years ahead. And for anyone building, buying, or using AI, there has never been a better time to start trusting your gut — and building your own vibe test.
TLDR: Benchmarks alone can no longer capture what makes an AI great to actually use. OpenAI co-founder Karpathy's search for the next AI "vibe test" — inspired by playful prompts involving unicorns, pelicans, and Middle-earth — points to a future where personality, imagination, and human feel become core measures of AI quality. Businesses should start building their own "vibe suites" today, blending playful prompts, diverse testers, and gut-feel feedback with traditional metrics, because the AI models that win the future will be the ones people genuinely enjoy talking to.