Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

EXIMO: A Three-Stage Method to Efficiently Fine-Tune VLA Robot Policies

AI By Crimson AI Hugging Face Papers 22 August 2026 · 00:00 11 views
Share: X Telegram

Researchers propose EXIMO, a novel algorithm that combines VLM-guided exploration, imitation learning, and residual off-policy RL to fine-tune large vision-language-action (VLA) robot policies, significantly improving sample efficiency and final performance.

EXIMO: A Three-Stage Method to Efficiently Fine-Tune VLA Robot Policies

Key points

Fine-tuning large vision-language-action (VLA) models for new robotic tasks remains a major challenge. Traditional behavior cloning requires hundreds of hours of expensive teleoperation data, while reinforcement learning (RL) is often sample-inefficient, especially for long-horizon tasks. Moreover, RL with VLA models poses additional difficulties due to their size and architecture.

To address these issues, researchers introduce EXIMO, an efficient algorithm that fine-tunes VLA policies in three stages: explore, imitate, and optimize. In the explore phase, a vision-language model (VLM) acts as a planner, breaking down complex long-horizon problems into shorter, more manageable sub-tasks for the VLA. Together, the VLM and VLA collect an orchestrated dataset on new tasks.

During the imitate phase, the VLA is fine-tuned using this orchestrated data. Finally, in the optimize stage, residual off-policy reinforcement learning further refines the policy. The researchers ablated all three stages and found that EXIMO significantly outperforms existing approaches in both sample efficiency and final performance.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4