World models have become a key area in AI research, aiming to capture the structure and dynamics of physical environments. A new paper from Hugging Face introduces FactorJEPA, a novel approach designed for crowded and chaotic urban scenes typical of the Global South, a regime the authors call DENSEWORLD.
The researchers argue that existing JEPA (Joint Embedding Predictive Architectures) models struggle in these environments, which feature soft spatial boundaries, extreme agent heterogeneity, persistent occlusion, and rapid social negotiation under mixed traffic. To address this, they created DENSEWORLD-115k, the first large-scale dataset of its kind, comprising 1,000 hours of drive-through, walk-through, and aerial video across 22 cities.
FactorJEPA replaces the monolithic future-latent predictor of V-JEPA with three structured subspaces: layout, agents, and interactions. A visibility gate ensures partially observed agents are downweighted rather than discarded, while cross-channel penalties discourage shortcuts. This design improves future-latent accuracy, intervention-sensitive prediction, and robustness to reduced visual evidence.
In experiments, FactorJEPA outperformed strong baselines like LoRA and full fine-tuning, with advantages of up to 33.2×, 13.9×, and 43.3× in paired confidence intervals for key metrics. The method's rankings were consistent across 1B and 2B V-JEPA backbones, with Spearman correlations between 0.895 and 0.978.
The dataset and checkpoints are publicly released, enabling further research in this underexplored regime.