Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

FactorJEPA: A New World Model for Crowded Global South Urban Scenes

AI By Crimson AI Hugging Face Papers 8 August 2026 · 00:00 19 views
Share: X Telegram

Hugging Face researchers introduce FactorJEPA, a world model that decomposes future prediction into layout, agent, and interaction channels, outperforming existing JEPA methods on a new dense urban dataset.

FactorJEPA: A New World Model for Crowded Global South Urban Scenes

Key points

World models have become a key area in AI research, aiming to capture the structure and dynamics of physical environments. A new paper from Hugging Face introduces FactorJEPA, a novel approach designed for crowded and chaotic urban scenes typical of the Global South, a regime the authors call DENSEWORLD.

The researchers argue that existing JEPA (Joint Embedding Predictive Architectures) models struggle in these environments, which feature soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. To address this, they created DENSEWORLD-115k, the first large-scale dataset of its kind, comprising 1,000 hours of drive-through, walk-through, and aerial video across 22 cities.

FactorJEPA replaces the monolithic future-latent predictor of V-JEPA with three structured subspaces: layout, agents, and interactions. A visibility gate ensures partially observed agents are downweighted rather than discarded, while cross-channel penalties discourage shortcuts. This design improves future-latent accuracy, intervention-sensitive prediction, and robustness to reduced visual evidence.

In experiments, FactorJEPA outperformed strong baselines like LoRA and full fine-tuning, with advantages of up to 33.2×, 13.9×, and 43.3× in paired confidence intervals for key metrics. The method's rankings were consistent across 1B and 2B V-JEPA backbones, with Spearman correlations between 0.895 and 0.978.

The dataset and checkpoints are publicly released, enabling further research in this underexplored regime.

MetricImprovement (vs. strongest baseline)
Future-frame L133.2× paired-CI
Causal L113.9× paired-CI
Mask-ratio slope43.3× paired-CI
Motion cosine20.0× paired-CI
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1