Hugging Face Trending Papers

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

Read the original on Hugging Face Trending Papers →

Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact. We present TextHOI-3D, a staged framework that uses generated multi-view observations as an explicit interface between text-conditioned visual generation and geometry-aware hand-object recovery.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jun 11

TextHOI-3D: Text-to-3D Hand-Object Interaction via Discrete Multi-View Generation and Joint Mesh Optimization

arXiv:2606. 11805v1 Announce Type: cross Abstract: Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact.

By Zixiong Hao, Zhencun Jiang
arXiv Computer Vision
Sep 21

VGGT-CAD: Reconstructing Parametric CAD 3D Model with Geometric Grounding

VGGT-CAD is a geometry‑aware framework that reconstructs parametric CAD 3D models from single and multi‑view images. It incorporates pretrained 3D geometric priors by encoding camera parameters as condition tokens and jointly modeling them with image tokens. The method introduces a variable‑view cross‑view context aggregation module and a training‑free geometry‑aware view selection strategy, and decodes the learned representation into CAD command sequences using a non‑autoregressive decoder. Additionally, VideoCAD, a large‑scale multi‑view video benchmark derived from existing CAD data, is presented to evaluate the approach.

By Chunan Yu, Tianrun Chen, Fu Shen, Cheng Chen, Lanyun Zhu, Yang Yang
arXiv Computer Vision
Sep 23

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

The paper introduces Segment‑Snap, a method that jointly models movable parts, their motion, and the regions they operate on in 3D scenes. It uses learned predictors to identify part surfaces and handles, a geometric decoder to constrain motion with planar and upright priors, and a joint part‑and‑handle predictor to refine motion classes. Experiments on Articulate3D show that handle guidance boosts motion‑gated AP from 13.74% to 40.98%, while additional handle candidates and part‑based corrections further improve performance.

By Hanyang Kong, Xingyi Yang