Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Research: CSCR Reallocates Token Credit to Improve Long-CoT Reasoning

AI By Crimson AI Hugging Face Papers 3 August 2026 · 00:00 9 views
Share: X Telegram

A new paper proposes Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens, consistently outperforming baselines on mathematical reasoning benchmarks.

Hugging Face Research: CSCR Reallocates Token Credit to Improve Long-CoT Reasoning

Key points

Reinforcement learning with verifiable rewards (RLVR) is a key driver for improving long chain-of-thought (CoT) reasoning in large language models. However, critic-free methods like GRPO uniformly distribute response-level advantages across all tokens, ignoring their unequal contributions to the final outcome. On-policy self-distillation (OPSD) offers dense supervision but implicitly assumes that likelihood shifts encode reliable answer-aligned information.

Researchers at Hugging Face tested this assumption by re-scoring trajectories under opposing outcome conditions. They found that most affected tokens shift in the same direction regardless of correctness, with few sign reversals and substantial overlap in optimization signals. Large shifts concentrate on substitutable surface-form tokens, while reasoning-critical tokens are less sensitive.

Based on these findings, they propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve the original credit budget and verifier-determined direction. On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms GRPO with the same number of policy updates.

Targeted ablations confirm that privilege-induced directions are unreliable, moderate downweighting is most effective, and stronger modulation destabilizes optimization. The method also improves over self-distillation baselines across multiple model scales.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1