Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

WorldWeaver: Streaming Multi-Agent Diffusion Model with Cross-Agent State Registers

AI By Crimson AI Hugging Face Papers 25 July 2026 · 00:00 9 views
Share: X Telegram

Hugging Face researchers introduce WorldWeaver, a streaming multi-agent video diffusion model that uses learnable world state registers to maintain shared information across agents and views, improving consistency in multi-agent Minecraft video generation.

WorldWeaver: Streaming Multi-Agent Diffusion Model with Cross-Agent State Registers

Key points

Hugging Face researchers have published a paper on WorldWeaver (W^2), a streaming multi-agent video diffusion model designed to maintain consistent world states across multiple agents and viewpoints. The model addresses a key limitation of existing autoregressive video diffusion pipelines, which carry forward observation history as conditioning context but struggle to maintain shared state in multi-agent and multi-view settings.

WorldWeaver introduces cross-agent world state registers: learnable tokens that store shared world information, track individual agent status, and are dynamically updated after each generated chunk. These registers are grounded with supervision signals spanning individual agent status, global state views (including bird's-eye views), and scene text.

The architecture incorporates a Mixture-of-Transformers design with separate weights for world state modeling and visual frame modeling. Extensive experiments in two-agent Minecraft video generation demonstrate that explicit world-state modeling improves logical consistency and generation quality.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1