Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Researchers Tackle Trajectory Anchoring Bias in Autonomous Driving VLMs with DEFT-RLVR

AI By Crimson AI Hugging Face Papers 4 August 2026 · 00:00 10 views
Share: X Telegram

A new paper from Hugging Face researchers identifies a trajectory anchoring bias in autonomous driving VLMs caused by ground-truth-conditioned chain-of-thought supervision, and proposes AD-MCQ and DEFT-RLVR to make planning verifiable and causally faithful.

Hugging Face Researchers Tackle Trajectory Anchoring Bias in Autonomous Driving VLMs with DEFT-RLVR

Key points

Researchers at Hugging Face have published a new paper that exposes a critical flaw in how autonomous driving vision-language models (VLMs) are trained. The study, titled "Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs," identifies a trajectory anchoring bias that arises when chain-of-thought (CoT) supervision is conditioned on the logged ground-truth (GT) future trajectory.

The authors demonstrate that when teacher models are shown the actual future trajectory during training, they tend to rationalize the revealed outcome rather than infer decisions from scene evidence. This leads to less causally faithful reasoning and significantly more severe hallucinations, particularly in causally challenging driving scenarios.

To address this, the team introduces Autonomous-Driving Multiple-Choice Question (AD-MCQ), which reformulates planning as a selection among explicit trajectory candidates, avoiding the need for open-ended trajectory synthesis. Building on this, they propose DEFT-RLVR (Deferred Exposure of Future Trajectories for RLVR), a reinforcement learning framework that transforms future trajectories from pre-decision anchors into post-decision verification targets.

Experimental results show that DEFT-RLVR improves AD reasoning while preserving or even enhancing general visual capabilities. Because AD-MCQ operates entirely within the VLM and allows difficulty to be controlled through candidate construction, it offers a flexible, scalable, and extensible foundation for future research on verifiable AD reasoning.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1