Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate.
Researchers propose CAST (Credit Assignment from Solver Teachers), which converts changes in a game solver's state value into solver advantages and injects them into RLVR as turn-level signals. The key insight is that changes in a game solver's state value reveal whether an action advances the state toward success.
Under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. This yields a form of on-policy, logit-free distillation that needs only one scalar per action.
Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. The code is available at https://github.com/Wloner0809/CAST.