DrivingBench puts four frontier models behind a real Toyota Corolla, with only GPT-6 Astra finishing the course

DrivingBench puts four frontier models behind a real Toyota Corolla, with only GPT-6 Astra finishing the course

N
News Editor
2026-09-30 09:36:27
Three engineers in the San Francisco Bay Area connected four general-purpose large language models to a real 2022 Toyota Corolla and tested them on a low-speed cone course in a parking lot. The setup used comma.ai’s comma four device, openpilot, and an MCP-based tool layer that let the models observe camera feeds, issue motion commands, and trigger an emergency stop. Results released on Sept. 21 showed that only OpenAI’s GPT-6 Astra completed the roughly 130-meter route, finishing on its second attempt in 5 minutes and 22 seconds. The report says the models did not reliably agree to drive at first. Some, especially GPT-6 Astra, sometimes refused to control a physical vehicle even after the prompt specified a cleared parking lot, low speed limits, and a human ready to brake. The team said performance became much more stable after renaming the MCP server to "DrivingBench Sandbox" and updating the prompt. Claude Fable 5.1 reached 45% at best, Grok 4.6 reached 11%, and GPT-5.6 Sol stopped at 6% in all three attempts. According to the report, the main failure mode was perception: reading the cone layout correctly from camera images. The team said the experiment was not meant to show that connecting ChatGPT-style systems to personal cars is a practical transportation setup, and plans to expand the benchmark with more runs, more models, and harder tracks.

DrivingBench, a project built by three engineers in the San Francisco Bay Area, connected four general-purpose large language models to a real 2022 Toyota Corolla and tested them on a low-speed course laid out with small traffic cones in a parking lot. The task was simple on paper: follow the route and park inside a blue cone-marked space at the end. Results released on Sept. 21 showed that only OpenAI’s GPT-6 Astra completed the full run.

The models did not reliably agree to drive until the MCP server was renamed

According to the DrivingBench report, the models were not consistently willing to control a real car at the start of testing. Some of them, especially GPT-6 Astra, sometimes refused on safety grounds. The prompt already stated that the parking lot had been fully cleared, that the test would run at low speed, and that a human was standing by with a foot over the brake, but refusals still happened.

The team said the behavior became much more stable after it renamed the MCP server that linked the models to the vehicle as "DrivingBench Sandbox" and updated the prompt. In the report, that naming change is described as the best-performing setup, even if the reason was unclear. Aditya Ramabadran told 404 Media that the group spent several hours revising prompts and labels before GPT-6 Astra would drive consistently. He also said that after switching to that version of the prompt, the model "almost never refuses anymore."

The team consists of Aditya Ramabadran, Tobias Gessler, and Simon Mahns. The three are colleagues at AI math startup Axiom Math.

How the vehicle was connected to the models

On the vehicle side, the team used comma.ai’s comma four, a windshield-mounted device that works with the open-source driver assistance software openpilot. Together, they can control steering and speed. The engineers added their own control software on top of openpilot, then handed operation to the models through MCP, short for Model Context Protocol.

DrivingBench exposed three tools to the models:

  • observe, which returned camera images and vehicle speed;
  • stop_now, which triggered immediate braking;
  • set_motion, which set direction, steering magnitude, speed, and duration.

The four models in the test were OpenAI’s GPT-6 Astra and GPT-5.6 Sol, Anthropic’s Claude Fable 5.1, and Grok 4.6. GPT-6 Astra and GPT-5.6 Sol ran in Codex, Claude Fable 5.1 ran in Claude Code, and Grok 4.6 ran in Cursor. All four were configured with medium reasoning intensity.

Only GPT-6 Astra finished the roughly 130-meter course

The course was set up in a large Bay Area parking lot and measured about 130 meters. It included a left turn, a straight, a gentle bend, and two right turns. To count as a finish, the car had to enter a parking space marked by blue cones. Each model got up to three attempts, all within the same conversation. After each failed run, the model was asked to review what went wrong before trying again.

Progress was measured by how far the car advanced along the centerline of the course. Any portion more than 4 meters away from that centerline did not count.

GPT-6 Astra reached 49% on its first attempt, then finished the course on its second in 5 minutes and 22 seconds. GPS measured the trip at 134.7 meters. That works out to an average speed of about 1.5 km/h, slower than a typical walking pace. The run used about 6.6 million tokens and, at list pricing, cost about $7.74.

Claude Fable 5.1 posted its best result on the third attempt, reaching 45%. The report said that was roughly comparable to GPT-6 Astra’s first run, or about halfway through the course. Its other attempts did not make it past the first turn. Grok 4.6 reached 11% at best, while GPT-5.6 Sol stopped at 6% in all three tries. GPT-6 Astra, the model singled out in the report as one that would sometimes refuse to drive, was also the only one to finish.

The main problem was perception

The report says the main failure mode was perception, meaning how the models interpreted the environment from camera images. At the start of the course, a diagonal line of cones forced the models to determine which side of that line defined the lane. After its second failed attempt, Claude Fable 5.1 wrote in its review: "I picked the wrong side of the boundary again. That diagonal cone line is the left edge of the lane, not the right edge." GPT-5.6 Sol, in its own review, said it had mistakenly treated cone colors as left-right boundary cues even though the prompt had already stated that the cones came in multiple colors.

The report says both GPT-6 Astra and Claude Fable 5.1 showed signs of learning within the conversation. After its first failed run, GPT-6 Astra wrote that it had decided too early that the car was aligned and had increased speed to 1.5 meters per second. It then noted that near turns it should slow to 0.5 to 0.8 meters per second. On the second run, its speed stayed below 0.8 meters per second, about 2.9 km/h, and 20 of its 24 commands set steering to the 100% limit.

Claude Fable 5.1 initially used steering values between 60% and 100%, then concluded in its review that the range was too aggressive and changed its default to 30%. On the third attempt, it explicitly noted that the diagonal cone line should remain on the left side and successfully cleared the long straight, but it did not leave enough room for the next right turn.

Low-speed parking lot test with a human ready to brake

The experiment took place entirely in an empty parking lot and never entered public roads. The software limited speed to 0.5 to 3.5 meters per second, or about 1.8 to 12.6 km/h. If speed exceeded 6 meters per second, about 21.6 km/h, the system would cancel the action and disengage control.

Each attempt was started by a human in the driver’s seat pressing the RES button on the steering wheel. Under openpilot’s rules, someone must press that button before the car can move from a standstill. The driver kept a foot over the brake throughout the test and intervened if the vehicle left the course or approached an obstacle or curb.

One detail in the report stands out: the car did not stop while the model was thinking. The previous command kept running until its time limit expired. A system diagram in the report shows model thinking times of about 2 to 30 seconds per step. Ramabadran told 404 Media that if the car is moving at 1 meter per second and the model thinks for 10 seconds, the vehicle has already traveled 10 meters.

The report says GPT-6 Astra checked the camera feed about once every 5 to 6 seconds. Claude Fable 5.1’s second attempt lasted 190 seconds, but the car moved for only 31 seconds, with most of the remaining time spent waiting for inference. Only GPT-6 Astra and GPT-5.6 Sol sent a new command before the previous one had fully completed.

NHTSA opened a preliminary probe into comma.ai on the same day

The comma four device used in the test is also part of a U.S. National Highway Traffic Safety Administration, or NHTSA, investigation. On Sept. 21, NHTSA opened a preliminary probe into comma.ai covering an estimated 30,000 devices, including comma three, comma 3X, comma four, and the openpilot software.

The probe stems from five crashes in which vehicles equipped with comma devices struck stopped or slow-moving vehicles in the same lane. Two of those crashes resulted in three deaths, and 11 people were injured. The report says that investigation concerns road incidents and is unrelated to the DrivingBench parking lot experiment.

The team says it was not trying to prove this is a practical way to drive

Ramabadran told 404 Media that the team was not trying to show that connecting ChatGPT to a personal car is a workable transportation setup. The next version of the benchmark will repeat tests more times for each model and compare different reasoning settings. The team also plans to add more models and use longer or more difficult tracks.

The report concludes that the fact a model could succeed in this test shows there is still urgent work to do on safety, alignment, and evaluation. The team said it plans to share more on that soon.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
200

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.