GPT-6 Astra drives a real Toyota Corolla: inside the DrivingBench test
Three engineers let GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol and Grok 4.6 drive a real Toyota Corolla around a cone course. Only Astra finished.

Three AI engineers wanted lunch at In-N-Out Burger, and they decided not to drive there themselves. Sitting in a 2024 Toyota Corolla near the drive-thru lane of a Bay Area restaurant, they opened a laptop and asked OpenAI's GPT-6 Astra to take the wheel. According to a report by Will Knight in WIRED, published on October 7, 2026, the model slowly but surely steered the car up to the take-out window so they could collect their food. A safety driver kept one foot above the brake the whole time.
The stunt was more than a lunch run. The same trio, Aditya Ramabadran, Simon Mahns and Tobias Gessler, who work at a startup called Axiom, turned the idea into a proper test. They built a benchmark called DrivingBench and published a paper on arXiv that puts four general purpose AI models behind the wheel of the same real car on a cone course in a parking lot. Only one model finished: GPT-6 Astra. Below is what they did, how the models performed, why the models first refused to drive, and what this does and does not tell us about AI in the physical world.

The three things you need to know
- The setup. A general purpose chat model, not a self driving system, saw camera frames from a Toyota Corolla and sent steering and speed commands to the car through a small set of tools. The tests ran at low speed in a parking lot, with a safety driver ready to brake.
- The results. On the DrivingBench cone course, GPT-6 Astra was the only model to finish, on its second attempt. According to WIRED, Claude Fable 5.1 made it 45 percent of the way around and Grok made it just 11 percent. The paper says no other attempt passed 50 percent of the course.
- The catch. The models first refused to control a physical car. Only careful prompting, and the way the task and the tools were framed, got them to drive at all. The authors say they release their harness, prompts, course map and traces so others can check the work.
What the engineers actually built
Self driving cars are not new. Companies like Waymo and Tesla use systems that are specifically trained and engineered for driving, with years of data and dedicated software. What makes this experiment different is that the driver was a model designed to output text, code and the occasional image. WIRED describes the setup this way: the engineers linked a chat interface to a server, which was connected to several cameras mounted on the windscreen and to the car's power steering system. The model got no prior coaching for the task. It was asked to drive on the fly.
The DrivingBench paper describes the same idea in more formal terms. Through three tools, the models see camera frames from the Corolla and directly command its steering and velocity around a parking lot cone course at low speeds. Two details make the task harder than it sounds. First, the car may keep moving while the model is still thinking. Second, a new command replaces the one that is currently running. That means inference latency, the time a model needs to produce an answer, is part of the challenge. A model that thinks slowly is effectively driving with a delay.
The authors say this tests whether a model can observe, act, monitor, recover from mistakes and finish a long task under those constraints. That is a very different skill from answering a question in a chat window, where nothing in the world moves while the model works.
How the four models performed
The paper benchmarks four models: GPT-6 Astra, Claude Fable 5.1, GPT-5.6 Sol and Grok 4.6. Each model ran in what the authors call a vendor native harness, meaning tools such as Codex, Claude Code and Cursor, and each got up to three attempts within one conversation. That last point matters, because a model could learn from its earlier failed attempts while the context was still there.
The headline result is simple. GPT-6 Astra is the only model that finished the course, and it did so on its second attempt. WIRED adds that it completed the course very slowly. No other attempt by any model passed 50 percent of the course. In WIRED's numbers, Claude Fable 5.1 made it 45 percent of the way around, and Grok made it just 11 percent.

The paper also reports that two of the four models improved materially across attempts with retained context. Ramabadran told WIRED that the latest models seemed able to adjust to errors. "The models really did seem to be adjusting or in-context learning based on their mistakes and learning how to better navigate the controls," he said. In other words, the models did not just follow a fixed plan. They appeared to adapt to how the car responded and got better at controlling it in real time.
It is worth being precise about what was measured. This was a simple course in a parking lot, at low speed, with a human ready to stop the car. It was not a road test, not traffic, and not a claim that any of these models can drive safely in the real world. The authors themselves frame it as a first benchmark, and WIRED's summary is that the models still have a long way to go before they could pass a driving test.
Why the models refused at first
One of the most interesting parts of the story is that the models did not want to drive. According to WIRED, the engineers first tested Grok, then the latest models from OpenAI and Anthropic. At first, the models refused to take control of the vehicle. They replied with lines like "I can help interpret road images, but I can't issue motion commands to a physical car."
With careful prompting, the models could be coaxed into going further. The DrivingBench paper treats this as a research finding in its own right. The authors detail the design principles behind their action interface and show how the format of the tool output and the framing of the task combined to determine whether the models would drive at all or refuse.

This cuts both ways. On one hand, it is reassuring that frontier models are cautious by default when asked to move a two ton machine. On the other hand, the experiment suggests that this caution can depend heavily on wording and framing. If the same model refuses one request and accepts a nearly identical one that is presented differently, then the safety behaviour is less a firm rule and more a reaction to how the task looks. That is a useful lesson for anyone building agents that control real hardware, from robots to cars to industrial equipment.
Where the driving skill might come from
None of the engineers believe the big AI labs are secretly training their models to drive cars. They told WIRED that the vehicular talent likely appeared as a side effect of training focused on 3D reasoning. The idea for the project came from noticing that models like Astra could build complex 3D simulations, which made the trio wonder whether that ability might translate into navigating the real world.
"This could be like an emergent capability of just scaling up the multimodality of the model," Mahns told WIRED, referring to inputs like images, video and 3D models that are now used to train models. Ramabadran added a lighter explanation for the idea itself: "We all live in the Bay Area, and we're exposed to Teslas and Waymos, so maybe that played a role."
The broader trend is real. WIRED notes that today's smartest models are very good at answering complex questions and at some virtual tasks, but they tend to struggle once they leave the world of computers and the internet. Physical reasoning is still seen as a frontier. Startups are forming around it. Andrew Dai, a former Google DeepMind researcher who now runs a company called Elorian AI, told WIRED that better visual reasoning will open up many new applications and called it "pretty essential for robotics." Elorian and Scale AI also recently developed a benchmark called Humanity's Sixth Sense, which measures how well models understand physical scenes.
What this means, and what it does not
It would be easy to turn this story into a headline like "chatbots can drive now." That is not what the evidence shows. A few things are true at the same time.
A text model really controlled a real car. That is notable. The model was not a dedicated driving system, it had no special training for the task, and it still managed to finish a cone course once and to reach a drive-thru window.
The skill is very limited. One model out of four finished the course, slowly, on its second try, in an empty parking lot at low speed. The others did not get past the halfway mark. Real roads add other drivers, pedestrians, weather, speed and rules that this test did not include.
Latency is a real problem. Because the car keeps moving while the model thinks, a slow answer is a dangerous answer. Large models that reason for several seconds are a poor fit for anything that needs split second reactions. This is one reason purpose built driving systems look very different from chat models.
Safety guardrails depend on framing. The models refused at first, and the paper shows that the tool output format and the framing of the task decided whether they would drive. That is worth watching closely as more AI agents are connected to the physical world.
The work is open to checking. The authors say they release their harness, prompts, course map and traces with video and telemetry for reproducibility. That makes DrivingBench something other researchers can rerun, criticise and extend, rather than a one off viral clip.
As one of the engineers joked during the lunchtime ride, "Maybe AGI is here after all." The more sober reading is that general purpose models are starting to build a rudimentary but useful understanding of the physical world, and that this understanding is still far from what it takes to drive. Experiments like this one help show exactly where that gap is.
Sources
- WIRED, Will Knight: These Researchers Made AI Drive a Toyota Corolla to Get In-N-Out (October 7, 2026)
- arXiv: DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?
Would you let a chatbot drive your car? Tell us on Reddit in r/djcroman or on X at @djcroman, and follow djcroman for your daily AI news.
Source: wired.com