Everyone has that one piece of furniture. The shelf that wobbles. The bookcase with the back panel nailed on upside down. The desk with one leg that never quite sat flush. You followed the instructions, mostly, and somehow it still came out wrong.
Now imagine pointing your phone at the half-built mess and getting a straight answer: "You flipped the side panel in step four. That's why the holes don't line up."
That is the promise behind OpenAI's GPT-6 Astra, announced in late September 2026. The headline capability is simple to say and hard to build: the model can look at what you're doing with your hands, compare it to what the instructions actually said, and pinpoint the exact step where things went sideways.
It sounds like a party trick. It isn't. This is a real shift in what AI can do, and it points to where the whole industry is heading next.
For years, the way we measured AI progress was simple. Could it answer questions? Could it write code? Could it pass an exam? Those are all text problems. The AI reads words, thinks in words, and answers in words.
Assembling a shelf is not a text problem. It's a physical one. It involves:
A model that can catch your step-four mistake while you're on step twelve is doing something fundamentally different from a chatbot. It is tracking a physical process over time and reasoning about cause and effect in the real world.
That is the leap. Not the shelf. The shelf is just the demo everyone can understand.
Here's the part business leaders should focus on. Until now, most AI tools waited for you to ask something. You typed a question. It typed back. The AI was a responder.
GPT-6 Astra, based on what this capability implies, is moving into a different role. It watches, it remembers what you did earlier, and it speaks up when something has gone wrong, even if you didn't think to ask.
Think about what that unlocks. Anywhere a human follows a process and can make a quiet mistake, an AI supervisor could sit alongside them:
The common thread: AI stops being a tool you consult and starts being a second set of eyes that never gets tired.
Language models got very good, very fast. But language is a compressed, tidy version of reality. Words are already organised. A sentence has a subject and a verb. Meaning is mostly explicit.
The physical world has none of that courtesy. A camera feed is a firehose of pixels with no labels. To make sense of it, a model has to build an internal picture of what's happening, what objects exist, where they are, how they relate, and how they're changing.
Getting that right requires several skills working together:
Each of those has existed in some form. Combining them into one system that works reliably is the hard part. That combination is what "spatial and process intelligence" means, and it is where the next round of competition in AI will be fought.
You don't need to build furniture to feel the impact. If your business involves people following procedures, and almost every business does, this is worth paying attention to.
Most companies train people by pairing them with an experienced colleague. That's expensive and slow. An AI that watches and corrects in real time could compress months of learning into weeks, and free up your best people to do actual work.
Today, quality checks happen at the end. You find the defect after the value has been added, which means rework, waste, and delays. Real-time process watching catches problems at the moment they occur, when fixing them is cheap.
Every company has procedure manuals nobody reads. When an AI can compare what's actually happening to what the manual says, the manual becomes live again. And when the AI keeps flagging the same step, you've learned something about your process, not just your people.
The retiring senior technician holds forty years of knowledge. Historically, that knowledge leaves with them. A system that can watch them work and coach others is a way to keep some of it in the building.
It would be easy to get carried away. A few things deserve real scrutiny.
Accuracy under pressure. A model that's right 95% of the time is helpful. In a factory, a hospital, or on a job site, the other 5% matters a great deal. Systems like this will need clear limits on what they're allowed to claim.
Who is accountable. If an AI misses a mistake and something goes wrong, who answers for it? Companies adopting this technology will need to decide that before deployment, not after.
Privacy and surveillance. A camera that watches you work is a camera that watches you work. There's a real difference between "this step was done wrong" and "you took eleven minutes longer than your colleague." The line has to be drawn deliberately, and preferably with employees in the room.
Over-reliance. If people stop thinking because something else is checking, skills can quietly erode. The best implementations will coach rather than replace judgement.
You don't need a budget line for this today. You do need to start thinking about it now.
Every major wave of AI has moved the technology one step closer to the messy, physical, real-time world we actually live in.
First, AI learned to handle text. Then it learned to handle images and sound. Now it's learning to follow a process unfolding in time, and to notice when that process goes off the rails.
That's a bigger deal than a flat-pack shelf. It means AI is graduating from something you talk to, into something that works alongside you and pays attention on your behalf. The IKEA demo is memorable because everyone has lived it. But the same capability, pointed at a factory line, an operating theatre, or a wind turbine in the North Sea, is where the real value sits.
The practical takeaway for leaders is straightforward. Start asking now which of your processes involve people doing physical steps in a fixed order, because those are the processes about to get a very patient, very observant partner.
And for everyone else? The next time you're three hours into a build and nothing lines up, help may finally be on the way.