Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. However, existing methods for articulated object reconstruction rely on explicitly observable motion from multiple articulation states, limiting their applicability.
In a new paper, researchers introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration. This inherently ill-posed setting compensates for the absence of motion cues by leveraging geometry, semantics, and motion priors.
The framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, the method uses a video diffusion model to synthesize articulation hypotheses and validates them through geometric consistency.
The approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines. It recovers part geometry and joint parameters from a single resting-state observation, exported directly as a simulation-ready URDF.