Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

V-RAE: Rethinking Video Latent Spaces for Generation

AI By Crimson AI Hugging Face Papers 19 August 2026 · 00:00 5 views
Share: X Telegram

Hugging Face researchers introduce V-RAE, a video representation autoencoder that builds compact generative latents from frozen vision representations, improving generation quality, convergence speed, and predictive modeling.

V-RAE: Rethinking Video Latent Spaces for Generation

Key points

Researchers at Hugging Face have introduced V-RAE (Video Representation Autoencoder), a novel approach to video latent spaces for generative modeling. The work addresses a key limitation of conventional video autoencoders: their latent spaces are optimized for pixel-level reconstruction, not for the semantic organization that generative models need.

V-RAE builds compact generative latents on top of frozen vision foundation model representations, such as DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal pooling module compresses frame-level representations while preserving semantic and motion information, and a video decoder reconstructs continuous motion from the compressed features.

In evaluations, V-RAE achieved a reconstruction FVD of 2.13 on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. Under matched generation settings, the best variant achieved gFVD scores of 117.86 on UCF101 and 19.16 on K600, while converging up to 6x faster.

The authors also introduce tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality than reconstruction FVD, highlighting that reconstruction quality alone is insufficient to characterize generative utility. Beyond generation, V-RAE improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched settings.

MetricV-RAE (best)Comparison
rFVD on K6002.13Outperforms all evaluated large-scale pretrained video VAEs
gFVD on UCF101117.86Matched generation settings
gFVD on K60019.16Matched generation settings
Convergence speedUp to 6x fasterCompared to conventional video VAEs
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4