CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction

Hao Xu1 Yilin Liu2 Yinqiao Wang1 Chi-Wing Fu1 Niloy J. Mitra2,3

1Chinese University of Hong Kong 2University College London 3Adobe Research

CHOIR teaser results

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. We present CHOIR, a Contact-aware HOI Reconstruction framework that uses contact as an explicit coupling signal between hands and objects. CHOIR initializes a coarse 4D HOI sequence from open-world visual priors, rectifies hand-object placement with a generative HOI spatial rectification module, and refines the result through contact-aware joint optimization. Experiments on controlled and challenging videos show that CHOIR improves object reconstruction, physical plausibility, and temporal consistency over state-of-the-art methods.

Method Overview

Stage 1 coarse 4D HOI initialization
From a monocular video, we first obtain 2D HOI cues and initialize 3D hand and object reconstructions to form isolated sequences.
Stage 2 generative HOI spatial rectification
A flow-matching generative HOI spatial rectification module rectifies hand-object relative placement before initial contact correspondence construction.
Stage 3 hierarchical contact-aware optimization
Contact-aware optimization refines the sequence across phases, yielding physically consistent 4D HOI sequences.

Our Representative Examples

1 / 0

Explore representative sequences with synchronized input video, reconstructed 4D geometry, and rest-pose contact predictions. Click on the above arrows to browse different representative examples.

Qualitative Comparison with SOTA Methods

Dataset click to switch
Toggle views click to switch
1/1 pages

More Qualitative Results of Our Method

Dataset click to switch
1/1 pages
Filter by click to switch
Camera
Motion
Contact

Bimanual Cases

Our method generalizes to challenging bimanual scenarios — two-hand two-object and two-hand one-object hand–object interaction. Drag the 3D viewer to explore the reconstruction from any angle.

Our code and training dataset will be released.