Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

PAWBench: Probabilistic Alignment in World Models Remains Elusive

AI By Crimson AI Hugging Face Papers 28 August 2026 · 00:00 2 views
Share: X Telegram

A new benchmark, PAWBench, reveals that current video generation models fail to match reference behavior distributions, highlighting a gap in probabilistic alignment for world modeling.

PAWBench: Probabilistic Alignment in World Models Remains Elusive

Key points

Recent advances in video generation have positioned these models as potential world models. However, a new study from Hugging Face researchers introduces a critical distinction: a true world model must not only generate plausible trajectories but also reproduce the distribution of possible behaviors under the same initial observation and action. This requirement, termed probabilistic alignment, is the focus of the new benchmark PAWBench.

The paper, titled "PAWBench: How Far Are We from Probabilistically Aligned World Modeling?", formalizes probabilistic alignment as a distributional criterion. To evaluate video generators as stochastic samplers, the authors introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over physical behaviors.

Testing across 50 scenarios and eleven current systems, the study finds that no model consistently matches the reference probabilities while recovering the full range of valid behaviors. This gap underscores a fundamental limitation in current world models.

The researchers also explore whether language prompts, initial noise sampling, or model training can reshape the predictive distribution, but the results suggest that these interventions are insufficient to achieve probabilistic alignment. The work aims to serve as a foundation for future efforts toward more reliable world modeling.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4