GPT-6 Astra Completes a Real Car Test in DrivingBench
Table of Contents
OpenAI’s GPT-6 Astra has completed a controlled real-car test conducted by the independent DrivingBench project, becoming the first model tested in the benchmark to finish its cone course. Astra guided a Toyota Corolla through the 134.7-metre course in 5 minutes and 22 seconds on its second attempt.
The experiment took place in a parking lot at very low speeds, with a human operator ready to intervene. The result demonstrates that a general-purpose AI model can perform a vehicle-control task under restricted conditions, but it does not establish that GPT-6 Astra is suitable for autonomous driving on public roads.
How the DrivingBench Test Worked
DrivingBench was created by researchers Tobias Gessler, Aditya Ramabadran and Simon Mahns to evaluate whether frontier AI models could control a physical vehicle. The benchmark gives models access to a Toyota Corolla’s steering, accelerator and brakes and measures their performance on a fixed cone course.
The test vehicle used a Comma 4 driver-assistance device running openpilot software. A laptop and phone were also part of the setup, while the AI model received camera observations and issued vehicle-control commands.
Astra’s first attempt reached 49% of the course before stopping. On its second attempt, it completed the full 134.7 metres in 5 minutes and 22 seconds. DrivingBench records the successful run at an average speed of about 0.94 miles per hour.
Other Models Could Not Finish
DrivingBench also tested Claude Fable 5.1, Grok 4.6 and GPT-5.6 Sol.
None of those models completed the course. Claude Fable 5.1 reached 45% on its best attempt, while Grok 4.6 reached 11% and GPT-5.6 Sol reached 6%.
The benchmark showed that perception was one of the difficulties for the models. The DrivingBench report noted that several attempts failed around the first corner, where the models had to determine which side of a diagonal line of cones represented the correct path.
These results are specific to this benchmark and should not be interpreted as a general ranking of the models’ overall driving capabilities.
The Test Came With a High AI Inference Cost
Astra’s successful run used about 6.6 million tokens and cost approximately $7.74 in inference fees, according to DrivingBench.
Because the vehicle travelled only 134.7 metres, that works out to approximately $92.47 per mile in inference cost for this particular benchmark run. This is not a projected operating cost for an autonomous vehicle, as it excludes hardware, insurance, supervision and other expenses.
The researchers also used a Comma 4 device costing around $999 as part of the test setup. The Register reported that a phone in the system connected to servers hosting GPT-6 Astra.
Processing Time Was a Major Limitation
The experiment also demonstrated the difficulty of using a general-purpose frontier model for time-sensitive physical tasks.
Astra used an observation tool that accessed the vehicle’s camera views approximately every five to six seconds during its successful attempt. This is very different from a conventional autonomous-driving system that continuously processes its surroundings and makes rapid decisions.
The latency was even more apparent in some other tests. The researchers reported that during one Claude Fable 5.1 attempt, the car moved for only 31 seconds during a 190-second run, with much of the remaining time spent waiting while the model processed its next action.
Safety Restrictions Remained in Place
The test was conducted under controlled conditions, with the vehicle moving at a very low speed and a human ready to apply the brakes.
The researchers also encountered cases in which models refused to control the physical car because of safety restrictions. According to the project report, they experimented with describing the environment as a simulation and eventually used the name “DrivingBench Sandbox” for the control server to reduce those refusals.
This detail is important because it shows that the experiment was not simply a case of connecting an AI model to a car and allowing it to operate independently.
Not a Demonstration of Self-Driving Readiness
The successful run is notable because GPT-6 Astra completed a physical driving task that other tested models did not. However, the benchmark was deliberately limited.
The course was short, the vehicle operated at extremely low speeds, the surroundings were controlled and a human was prepared to intervene. The model also relied on external computing and network connectivity.
DrivingBench researcher Aditya Ramabadran told The Register that using a frontier model directly for real driving is not currently practical. He cited the low-speed restrictions, model latency and the need for human supervision as important limitations.
The researchers suggested that a possible longer-term approach could involve training highly capable models and then distilling their capabilities into smaller, specialised systems capable of running on vehicle hardware.
What the Result Shows
The DrivingBench experiment provides a controlled demonstration of a general-purpose AI model interacting with a physical vehicle.
GPT-6 Astra was able to combine visual observations, planning and vehicle-control commands well enough to complete the benchmark course. At the same time, the experiment exposed significant challenges involving latency, inference costs, limited speed and the need for human oversight.
The result is therefore best understood as a proof of concept in a controlled environment, rather than evidence that GPT-6 Astra can safely drive on public roads.
FAQs
What is DrivingBench?
DrivingBench is a research benchmark that gives frontier AI models control of a Toyota Corolla’s steering, accelerator and brakes and evaluates them on a fixed cone course.
Did GPT-6 Astra drive a real car?
Yes. GPT-6 Astra controlled a real Toyota Corolla during the DrivingBench experiment and completed the 134.7-metre course on its second attempt.
How long did the successful attempt take?
The completed attempt took 5 minutes and 22 seconds.
Was a human involved?
Yes. The experiment was conducted with human supervision and safety measures, including very low vehicle speeds and readiness to intervene.
Is GPT-6 Astra ready for autonomous road driving?
The DrivingBench result does not demonstrate that. The test was conducted on a controlled course at very low speed with human supervision and external computing.
Conclusion
GPT-6 Astra’s DrivingBench result shows that a general-purpose AI model can control a real vehicle well enough to complete a simple course under carefully controlled conditions.
The experiment also highlights the substantial gap between completing a benchmark and operating a vehicle autonomously in everyday traffic. High inference costs, processing delays, network dependence and human supervision remain significant limitations.
For now, the result is best viewed as an early demonstration of AI-based physical control, rather than a production-ready autonomous-driving system.