Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Latent-to-4D: Direct 4D Generation from Video Diffusion Latents

AI By Crimson AI Hugging Face Papers 12 August 2026 · 00:00 11 views
Share: X Telegram

A new method, Latent-to-4D, enables reusable direct 4D generation from video diffusion latents, bypassing RGB and transferring across generators without retraining.

Latent-to-4D: Direct 4D Generation from Video Diffusion Latents

Key points

Researchers have introduced Latent-to-4D, a novel approach for generating dynamic 3D scenes (4D) directly from video diffusion latents. The method aligns a video latent with the token grid of a pretrained 4D decoder, refining it through frame-wise and global spatiotemporal attention. This bypasses the traditional RGB reconstruction step, avoiding distribution mismatch and error propagation.

Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a specific video generator to predict geometry directly. The former suffers from error accumulation, while the latter ties 4D prediction to a particular generator, often requiring retraining when the generator changes. Latent-to-4D instead leverages the shared variational autoencoder (VAE) of video models to create a reusable interface.

Trained on roughly 1,000 existing reconstruction clips, a single checkpoint of Latent-to-4D transfers unchanged across multiple video diffusion transformers within the same VAE family. This demonstrates significant flexibility and efficiency compared to prior approaches.

On the Text4D-200 and I4D-200 benchmarks, Latent-to-4D outperforms matched same-latent Wan+4RC cascades in projection-based DINO-F1 by 2.88–3.45 and 5.81 points, respectively. Human raters also preferred its outputs for geometry, temporal stability, and overall quality.

BenchmarkMetricLatent-to-4D vs Wan+4RC (improvement)
Text4D-200DINO-F1+2.88 to +3.45
I4D-200DINO-F1+5.81
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

0
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

0
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1