Research Papers
Predictive Divergence Masks: A New Direction for LLM Reinforcement Learning
Researchers propose predictive divergence masks to replace the ratio-based direction criterion in PPO-style RL for LLMs, improving...
ReOPD: Off-Environment On-Policy Distillation for Multi-Turn LLM Agents
Hugging Face researchers propose ReOPD, a method that reuses pre-collected teacher trajectories as replayed prefixes to achieve on...
Hugging Face Unveils Robostral Navigate: An 8B VLM for Scalable Robot Navigation Using Only a Single RGB Camera
Robostral Navigate is an 8B vision-language model that consumes only monocular RGB images to predict waypoints, achieving state-of...
Hugging Face Researchers Introduce Experience Distillation: Retaining 64.8% of In-Context Learning Gains Without Context
A new method called Experience Distillation allows agents to internalize interaction histories into model weights without addition...
Hugging Face Researchers Introduce Influence Matching for Dataset Distillation
Influence Matching (Inf-Match) aligns the final outcome of training by learning a compact synthetic set whose effect on converged...
GraphVid: Interactive Graph-Controllable Video Generation Achieves Superior Control with Structured Graphs
Hugging Face researchers introduce GraphVid, a graph-conditioned video generation model that uses structured interaction graphs fo...
Hugging Face Unveils FinanceComplexQA: A New Benchmark for Agentic Reasoning in Finance
FinanceComplexQA benchmarks agentic reasoning on industrial-grade financial documents, featuring 2,026 deep research tasks across...
OpenForgeRL: Open-Source Framework for Training Harness-Based AI Agents End-to-End
Hugging Face researchers introduce OpenForgeRL, an open-source framework that enables end-to-end reinforcement learning for AI age...
WorldWeaver: Streaming Multi-Agent Diffusion Model with Cross-Agent State Registers
Hugging Face researchers introduce WorldWeaver, a streaming multi-agent video diffusion model that uses learnable world state regi...
Hugging Face Researchers Propose Color Pass-Through: End-to-End Camera-Display Calibration
A new end-to-end learned framework treats camera and display as a coupled system, achieving over 2x improvement in color reproduct...
Hugging Face Researchers Propose Structured Dynamics Model to Disentangle Camera and Object Motion from Videos
A new self-supervised method, the Structured Dynamics Model (SDM), explicitly separates camera motion from object motion in videos...
LLMs Struggle to Track Evolving User Intent, New Study Finds
A new framework reveals that LLMs suffer significant performance drops when user intent evolves across conversation turns, exposin...