Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications in robotics training and virtual agents. Existing models remain limited: some focus on simple activities with static objects, while others restrict attention to hand-only manipulation.
Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of...
arXiv:2609.38466v1 Announce Type: new
Abstract: Text-conditioned full-body human-object interaction (HOI) generation requires synthesizing human motion and object trajectories that match the input te...
By Chuqiao Li, Xianghui Xie, Yong Cao, Andreas Geiger, Gerard Pons-Moll
arXiv:2609.36454v1 Announce Type: new
Abstract: We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanic...
By Wenliang Guo, Zhanbo Huang, Yu Kong
arXiv:2609.16683v1 Announce Type: cross
Abstract: Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and objec...
By Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu, Mingzhi Pei, Ruoqu Chen, Mengdi Xu
Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact. We present TextHOI-3D, a staged framework that uses generated multi-view observations as an explicit interface between text-conditioned visual generation and geometry-aware hand-object recovery.
Multi-finger dexterous manipulation relies on stable hand-object interactions, yet these interactions are partially observable in practice. Visual observations are often occluded by the hand, tactile...
arXiv:2608.13014v2 Announce Type: replace
Abstract: Understanding hand-object interaction from egocentric vision is essential for modeling how people physically engage with the surrounding world. Yet...
By Andela Ilic, Rachel Schuchert, Yijing Jiang, Christian Holz
arXiv:2606. 11805v1 Announce Type: cross Abstract: Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact.
By Zixiong Hao, Zhencun Jiang
The paper introduces a metric interaction framework for robotic manipulation that explicitly models object- and scene-level interactions in Cartesian space. It uses Interaction‑Centric Tokens (ICTs) to represent end‑effector trajectories relative to objects and a Metric Action Interaction Field (MAIF) to attend to scene point‑cloud features for geometry‑conditioned action corrections. Experiments show modest but consistent improvements across several benchmarks, including LIBERO, RoboTwin 2.0, and real‑world tasks.
By Lijie Wang, Zheng Lu, Yiming Wang, Heyang Yu, Kenghou Hoi, Bowen Hu, Di Cui, Tianyu Xin, Haoran Liao, Wanqi Zhong, Xingjie Fan, Yizhao Xu, Ziliang Wang, Fei Gao, Yiming Li
World models offer a promising route toward robot planning by enabling agents to imagine and verify the consequences of actions before execution. However, current video-based world models often struggle to capture the physical constraints that govern manipulation, particularly contact.
Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Hum...