Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections.
To address this, Hugging Face researchers introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. The model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision.
To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. The team further enhances target identification by co-training on a custom Refer-QA dataset.
On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4% on the single-target split, 73.3% on the distracted split, and 74.1% on the ambiguity tracking split — matching or even surpassing several multi-camera baselines on identification-heavy tasks.
Real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities. Code, data, and checkpoints will be released on GitHub at github.com/MedlarTea/referTrack.