DrivingBench Puts General-Purpose AI Behind The Wheel Of A Real Corolla
The interesting result is not that a language model can steer, but that safety training made it refuse a physical task it recognized as real, which is a benchmark finding about general-purpose AI in the world rather than about driving.
Reporting from 1 source: GIGAZINE.
Three Bay Area engineers ran an experiment called DrivingBench, connecting general-purpose AI models to a 2022 Toyota Corolla through the comma four device and openpilot software. The AI called three operations: check surroundings, move, stop. Told to drive a cone-lined reverse U-shaped course, several models refused once they recognized they were operating a real car, GPT-6 Astra most often. The team rewrote prompts to get runs completed.
The setup was deliberately unglamorous: a parking lot, cones, low speed, and a human in the driver's seat ready to brake. The AI saw cabin camera images and vehicle speed, then issued direction, steering, speed, and duration. The team's own framing is that DrivingBench is a benchmark for physical tasks, not a route to a self-driving product. What it measured first was reluctance. Models that would happily discuss driving balked at doing it, and the engineers spent hours on prompt and system-name wording before runs went through.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.