Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Study: On-Policy Delta Distillation Boosts Multilingual Math Reasoning

AI By Crimson AI Hugging Face Papers 7 August 2026 · 00:00 17 views
Share: X Telegram

A new paper from Hugging Face explores On-Policy Delta Distillation (OPD²) for math reasoning in English, Korean, and Japanese, showing consistent gains over standard OPD and a narrowing of the English-Korean gap.

Hugging Face Study: On-Policy Delta Distillation Boosts Multilingual Math Reasoning

Key points

On-Policy Distillation (OPD) is gaining traction as an alternative to reinforcement learning for post-training large language models, but its effectiveness in multilingual contexts has been largely unexplored. A new paper from Hugging Face investigates OPD and its advanced variant, On-Policy Delta Distillation (OPD²), specifically for mathematical reasoning in English, Korean, and Japanese.

OPD² improves upon OPD by using the probability gap between a post-trained teacher model and its base model as the learning signal, rather than relying solely on the teacher's outputs. In experiments with the Qwen3 model family, OPD² consistently outperformed the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrowed the performance gap between English and Korean.

The study also found that English-only OPD can boost performance for Korean and Japanese, but it often shifts the model's responses toward English, underscoring the importance of including multilingual data to preserve target-language responses. This highlights a key trade-off for practitioners aiming to improve reasoning in non-English languages.

The findings suggest that OPD² offers a promising path for efficient, multilingual post-training, potentially reducing the need for extensive reinforcement learning pipelines while maintaining or improving reasoning quality across languages.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1