arXiv AI

ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language Models

arXiv:2502. 08884v3 Announce Type: replace-cross Abstract: We present ShapeLib, the first method that uses the priors of Large Language Models (LLMs) to design libraries of programmatic 3D shape abstractions.

arXiv Computer Vision
Sep 23

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Fysiverse-3D-Vision is a unified vision‑language‑geometry framework that reconstructs executable 3D scenes from a single image. It separates spatial layout reasoning from asset synthesis, using a shared representation where spatial reasoning and geometric reconstruction reinforce each other. The model employs a Transformer that integrates textual supervision, semantic visual cues, and geometric representations, and includes an object‑conditioned layout module to predict object translation, rotation, and scale while maintaining physical consistency through collision‑aware optimization.

By Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang
arXiv AI
Jul 28

DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces

arXiv:2603. 05607v2 Announce Type: replace-cross Abstract: Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by small annotated datasets with explicit design histories or boundary representation (BRep) labels.

By Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias, Didier Stricker, Muhammad Zeshan Afzal, Jiankang Deng, Ismail Elezi
arXiv Computer Vision
Sep 11

Language-Augmented Semantic Priors for B-Spline Surface Fitting

The paper introduces LASP, a framework that uses large language models to generate structured B‑spline priors from procedural modeling histories. By translating design intent into rich textual descriptions, LASP provides semantic reasoning that guides conventional CAD solvers toward more accurate and coherent surface fitting. Experiments show that language‑driven priors outperform traditional machine learning approaches, establishing a new paradigm for language‑guided geometric optimization.

By Yunzhong Lou, Yusheng Luo, Jiahao Li, Yu Song, Xiangdong Zhou
arXiv Computer Vision
Sep 4

TokenMatch: 3D Mesh Correspondence Transformer with Curvature-Guided Tokenisation

TokenMatch is a transformer-based model that estimates 3D shape correspondences by adaptively tokenising meshes into curvature-guided patches. Trained only on the BeCoS partial-to-partial dataset, it generalises to full-shape matching without retraining, using self‑ and cross‑attention to learn patch‑ and point‑level relations. Evaluated on CP2P, PSMAL, BeCoS, FAUST, SCAPE, and SHREC'19, TokenMatch consistently outperforms existing methods in mean geodesic error and intersection‑over‑union while achieving sub‑second inference speeds.

By Adeela Islam, Zorah L\"ahner, Vittorio Murino, Vladislav Golyanik
arXiv AI
Jun 4

From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models

arXiv:2606. 04381v1 Announce Type: cross Abstract: Recent large language models (LLMs) often appear to exhibit spatial reasoning ability; however, this capability is largely \emph{symbolic}, arising from pattern matching over spatial language rather than true \emph{geometric} reasoning over space.

By Chen Chu, Bita Azarijoo, Li Xiong, Khurram Shafique, Cyrus Shahabi
arXiv Machine Learning
Sep 2

CADKnitter: Compositional CAD Generation from Text and Geometry Guidance

CADKnitter is a compositional CAD generation framework that uses geometric-guiding cues to steer diffusion sampling, enabling the creation of complementary CAD parts that satisfy both geometric constraints of an existing model and semantic constraints from a text prompt. The authors introduce KnitCAD, a dataset of over 310,000 CAD models paired with textual prompts and assembly metadata to support training and evaluation. Experiments show that CADKnitter outperforms state‑of‑the‑art baselines by a clear margin.

By Tri Le, Khang Nguyen, Baoru Huang, Tung D. Ta, Anh Nguyen
arXiv Computer Vision
Sep 3

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

The paper introduces a generalizable deformation learning framework that reconstructs 3D objects by deforming a category-level shape template to match a monocular observation. It employs a geometry-guided feature modeling mechanism to enrich foundation features with template topology, creating a geometry-aware representation that is explicitly correlated with the target observation for precise deformation. A view-adaptive feature aggregation module further bridges the gap between the fixed template and arbitrary target views by leveraging multi-view template features and camera poses, ensuring robust feature alignment across diverse viewpoints.

By Yiyao Ma, Kai Chen, Zhongxiang Zhou, Zhuheng Song, Dongsheng Xie, Zelong Tan, Rong Xiong, Qi Dou