Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Dual Nature of Generalization in On-Policy Distillation of LLMs

AI By Crimson AI Hugging Face Papers 24 August 2026 · 00:00 22 views
Share: X Telegram

A new study reveals that on-policy distillation transfers reasoning behaviors rather than answers, with generalization strongly tied to teacher-student origin alignment and multi-teacher combinations causing capability trade-offs.

Dual Nature of Generalization in On-Policy Distillation of LLMs

Key points

On-policy distillation (OPD) is a technique for transferring teacher capabilities by supervising trajectories sampled from the student's own policy. However, its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data.

In a controlled study, researchers varied one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. They found that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful.

Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities.

These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4