Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Researchers Propose Structured Dynamics Model to Disentangle Camera and Object Motion from Videos

AI By Crimson AI Hugging Face Papers 25 July 2026 · 00:00 14 views
Share: X Telegram

A new self-supervised method, the Structured Dynamics Model (SDM), explicitly separates camera motion from object motion in videos using frozen pretrained image features, outperforming baselines on the new ProbeMotion benchmark.

Hugging Face Researchers Propose Structured Dynamics Model to Disentangle Camera and Object Motion from Videos

Key points

Understanding motion in video remains a core challenge in visual learning, as frame-to-frame changes mix two distinct sources: camera motion and object motion. Decomposing these factors has been difficult due to their tight coupling in natural videos and the lack of separate supervision. Researchers from Hugging Face and academic partners introduce the Structured Dynamics Model (SDM), a self-supervised framework that recovers structured motion representations from frozen features of a pretrained image Vision Transformer (ViT).

SDM explicitly separates the dominant source of temporal change (typically camera motion) from residual dynamics (object motion) through future-feature prediction. Instead of using a single entangled latent or unstructured dense tokens, SDM learns a structured decomposition. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data.

The team evaluates SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision.

These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics. The project page is available at lukasknobel.github.io/projects/StructuredDynamics.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1