Results
Closed-loop results across the standard and safety benchmarks
We report closed-loop results following the order in the paper: overall driving quality on standard Bench2Drive, navigation-following capability on CARLA-F, navigation–scene conflict awareness on B2D-C and NavSim-C, and robustness to adversarial instructions on B2D-Adv. A strong language-conditioned driver should reach state-of-the-art driving quality and instruction following, keep infractions low when the navigation signal conflicts with the scene, and refuse instructions that are unsafe under any circumstance.
Standard Bench2Drive closed-loop benchmark. Nav. = navigation signal (CMD command, WP waypoint, Lan language). * trained with ×0.2 data scale. Higher is better for all metrics. SafeDriveVLA reaches state-of-the-art DS and SR.
| Method |
Nav. |
Expert |
DS ↑ |
SR (%) ↑ |
Efficiency ↑ |
Comfort. ↑ |
| UniAD | CMD | Think2Drive | 45.81 | 16.36 | 129.21 | 43.58 |
| VAD | CMD | Think2Drive | 42.35 | 15.00 | 157.94 | 46.01 |
| RecogDrive | CMD | Think2Drive | 71.36 | 45.45 | 138.18 | 17.45 |
| ORION | CMD | Think2Drive | 77.74 | 54.62 | 151.48 | 17.38 |
| MindDrive | CMD | Think2Drive | 78.04 | 55.09 | – | – |
| AutoVLA | Lan | PDM-Lite | 78.84 | 57.73 | 146.93 | 39.33 |
| DriveMoE | WP | Think2Drive | 74.22 | 48.64 | 175.96 | 15.31 |
| SimLingo* | WP | PDM-Lite | 81.48 | 53.66 | 246.01 | 42.33 |
| SafeDriveVLA* | WP | PDM-Lite | 83.64 | 59.82 | 260.67 | 52.48 |
CARLA-F navigation-following benchmark. Per-maneuver Navigation Compliance Rate (NCR, %); Speed is Speed Error (lower is better) and Avg. is mean NCR. * trained with ×0.2 data scale. From language alone, SafeDriveVLA leads all language-conditioned models and recovers lane-change NCR from 0%.
| Model |
Nav. |
Speed ↓ |
Turn left |
Turn right |
Go straight |
Left lane |
Right lane |
Lane follow |
Avg. ↑ |
| ORION | CMD | – | 100.0 | 100.0 | 94.1 | 60.6 | 58.6 | 86.2 | 87.8 |
| MindDrive | CMD | – | 100.0 | 92.2 | 94.1 | 33.3 | 37.9 | 70.7 | 78.8 |
| AutoMoT | CMD | – | 98.2 | 100.0 | 97.6 | 21.2 | 27.6 | 89.7 | 82.0 |
| SimLingo | Lan | 1.44 | 83.6 | 80.4 | 45.9 | 0.0 | 0.0 | 63.8 | 52.4 |
| SimLingo-IF | Lan | 1.31 | 89.1 | 82.3 | 49.4 | 0.0 | 0.0 | 69.0 | 55.6 |
| SimLingo-Safe | Lan | 1.67 | 92.7 | 84.3 | 50.0 | 0.0 | 0.0 | 62.1 | 55.6 |
| SafeDriveVLA* | Lan | 4.27 | 87.8 | 98.0 | 89.0 | 66.7 | 72.4 | 82.5 | 82.7 |
Bench2Drive navigation–scene conflict benchmark (B2D-C). Free-form language instruction is the sole navigation signal; we compare against SimLingo variants, the only baseline supporting language-conditioned driving. Higher DS / SR are better; lower infractions are better. * trained with ×0.2 data scale.
| Method |
DS ↑ |
SR (%) ↑ |
Collision ↓ |
Traffic Viol. ↓ |
Out of Route ↓ |
| SimLingo | 72.8 | 36.7 | 66 | 60 | 18 |
| SimLingo-IF | 56.3 | 12.7 | 166 | 57 | 32 |
| SimLingo-Safe | 72.8 | 38.0 | 60 | 61 | 11 |
| SafeDriveVLA* | 67.3 | 35.8 | 45 | 9 | 1 |
NavSim-C navigation–scene conflict benchmark. On the real-world NavSim split, existing VLAs degrade substantially when the injected command conflicts with the scene. Higher is better for all metrics; PDMS is the composite score.
| Method |
NC ↑ |
DAC ↑ |
TTC ↑ |
C. ↑ |
DDC ↑ |
EP ↑ |
PDMS ↑ |
| AutoVLA | 93.6 | 79.0 | 88.0 | 100.0 | 86.0 | 66.6 | 70.7 |
| ReCogDrive | 95.0 | 91.3 | 92.6 | 100.0 | 94.7 | 76.7 | 81.7 |
| CuriousVLA | 95.3 | 93.0 | 94.7 | 98.3 | 97.2 | 85.4 | 84.4 |
| DriveVLA-W0 | 96.3 | 87.0 | 91.7 | 100.0 | 97.0 | 71.0 | 76.9 |
B2D-Adv adversarial instruction benchmark. Driving Score on benign Bench2Drive routes (B2D) vs. adversarial instructions that explicitly try to override the safety prior (B2D-Adv); higher is better. SafeDriveVLA surpasses SimLingo by ≈70% under attack, and the ablation shows the mode token and world model each help.
| Method |
B2D DS ↑ |
B2D-Adv DS ↑ |
| SimLingo | 71.6 | 43.8 |
| baseline (no mode, no WM) | 76.3 | 71.4 |
| w/ mode | 81.6 | 77.5 |
| SafeDriveVLA* (w/ mode + WM) | 82.8 | 78.4 |