arXiv Computer Vision

MeshQuery: Agentic Seam Planning for UV Parametrization

arXiv AI
Aug 25

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.

By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang
Hugging Face Trending Papers
Jun 9

Dmsh: A Multi-Agent Reinforcement Learning Framework for All-Quad Mesh Generation

Generating high-quality meshes for arbitrary geometries remains a fundamental bottleneck in computational engineering, often demanding heuristic tuning and semi-manual workflows. In this paper, we introduce Dmsh, a first fully automated reinforcement learning pipeline that unifies geometric decomposition and quadrilateral mesh generation within a single learning-based framework.

arXiv Computer Vision
Sep 11

Language-Augmented Semantic Priors for B-Spline Surface Fitting

The paper introduces LASP, a framework that uses large language models to generate structured B‑spline priors from procedural modeling histories. By translating design intent into rich textual descriptions, LASP provides semantic reasoning that guides conventional CAD solvers toward more accurate and coherent surface fitting. Experiments show that language‑driven priors outperform traditional machine learning approaches, establishing a new paradigm for language‑guided geometric optimization.

By Yunzhong Lou, Yusheng Luo, Jiahao Li, Yu Song, Xiangdong Zhou
arXiv Computer Vision
Sep 30

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Beacon is a new agentic visual reasoning model that improves multimodal large language models (MLLMs) by better deciding when to use tools and how to use them. It introduces two key concepts—Mode Adaptiveness, which ensures tools are invoked only when necessary, and Tool Effect, which measures the net benefit of tool use— and trains the model with supervised fine‑tuning and reinforcement learning that rewards necessity-aware decisions and expands capability through expert hints. Across 13 benchmarks, Beacon outperforms other open‑source models, achieving the highest average score and the largest net tool‑gain on diagnostic tests.

By Qixun Wang, Yang Shi, Letian Cheng, Zhuoran Zhang, Yan He, Yuqi Tang, Qi Zhang, Xinlei Yu, Ruizhe Chen, Tianrun Xu, Yuanxing Zhang, Pengfei Wan, Haotian Wang, Xianghua Ying
arXiv AI
Aug 26

Design-to-Plan: A Large Language Model-Based Multi-Agent Framework for Manufacturing Process Planning from 3D CAD Models and 2D Engineering Drawings

Design-to-Plan is a large language model–based multi‑agent framework that automates end‑to‑end manufacturing process planning from 3D CAD models and 2D engineering drawings. The system uses an orchestrator to coordinate specialized agents for feature recognition, drawing analysis, context fusion, knowledge retrieval, process sequencing, tool selection, and report generation, integrating deterministic modules with LLM reasoning. Evaluation on 300 benchmark cases shows high success rates, strong tool selection accuracy, effective conflict detection, and reduced token usage, demonstrating the framework’s ability to produce consistent, traceable design‑to‑plan outputs.

By Muhammad Tayyab Khan, Lequn Chen, Wenhe Feng, Seung Ki Moon