arXiv:2606. 11805v1 Announce Type: cross Abstract: Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact.
By Zixiong Hao, Zhencun Jiang
arXiv:2608.28386v1 Announce Type: new
Abstract: Existing monocular full-body 3D human-object interaction (HOI) methods do not combine explicit finger-level grasp optimization with category-agnostic o...
By Semin Kim, Haechan Shin, Jongyoo Kim
Text driven hand object interaction (HOI) generation is gaining attention for immersive applications and robotics, yet producing physically plausible interactions remains challenging. Even when individual motions appear natural, small contact errors can cause conspicuous artifacts such as floating and interpenetration.
Estimating physical pressure from vision is essential for understanding contact-rich hand-object interaction. However, prior vision-based pressure estimation methods are largely limited to planar surfaces and single image input, making them difficult to apply to dynamic hand-object interaction with diverse objects.
VGGT-CAD is a geometry‑aware framework that reconstructs parametric CAD 3D models from single and multi‑view images. It incorporates pretrained 3D geometric priors by encoding camera parameters as condition tokens and jointly modeling them with image tokens. The method introduces a variable‑view cross‑view context aggregation module and a training‑free geometry‑aware view selection strategy, and decodes the learned representation into CAD command sequences using a non‑autoregressive decoder. Additionally, VideoCAD, a large‑scale multi‑view video benchmark derived from existing CAD data, is presented to evaluate the approach.
By Chunan Yu, Tianrun Chen, Fu Shen, Cheng Chen, Lanyun Zhu, Yang Yang
The paper introduces Segment‑Snap, a method that jointly models movable parts, their motion, and the regions they operate on in 3D scenes. It uses learned predictors to identify part surfaces and handles, a geometric decoder to constrain motion with planar and upright priors, and a joint part‑and‑handle predictor to refine motion classes. Experiments on Articulate3D show that handle guidance boosts motion‑gated AP from 13.74% to 40.98%, while additional handle candidates and part‑based corrections further improve performance.
By Hanyang Kong, Xingyi Yang