The paper introduces a novel framework that converts structural framing plans from PDF drawings into editable finite‑element model drafts without requiring task‑specific detector training. It combines a deterministic geometry extraction stage—estimating scale, recognizing five entity classes, and assembling a drafting grammar—with an agentic vision‑language refinement stage that proposes corrections, performs admission tests, and ensures fail‑closed transactions. Evaluation on a 100‑plan benchmark shows high accuracy, with scale within 0.1% and recall/precision values above 0.86 for all component types.
By Mohammad Talebi-Kalaleh, Qipei Mei
The paper investigates why reasoning‑augmented text‑to‑image models like GoT‑R1 sometimes fail on compositional prompts. By separating the explicit textual plan from the decoder, the authors show that the decoder faithfully executes the plan while the planner often writes incorrect spatial relations, especially for phrasing‑dependent cues. Editing or replacing the plan improves image quality without retraining, demonstrating the viability of modular planner‑decoder architectures.
By Ashritha Gonuguntla
We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records experiments over an extended period.
arXiv:2607. 25487v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets.
By Minhyeok Lee, Chiyoung Kim, Chanhoe Gu, Seongrok Kim, Sanghyuk Roy Choi, Donghwan Hwang, Donghun Ryu, Seokhyun Kim
arXiv:2607. 07984v1 Announce Type: new Abstract: Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task.
By Seokhoon Jeong, Mijung Kim, Taehwan Kim
In 2026, the SHROOM-Visions shared task was launched at the UncertaiNLP Workshop co‑located with EMNLP to address hallucinations in large vision‑language models. The task builds on the SHEEP dataset and asks participants to detect and classify fine‑grained hallucination spans in image‑conditioned text generation across four languages (Chinese, English, French, Italian) using a five‑class taxonomy. The competition attracted 27 teams and over 600 system submissions, with top systems achieving character‑level, label‑conditioned, and IoU scores of 0.58, 0.46, and 0.51 respectively, surpassing baselines by 30‑40 points.
By Ra\'ul V\'azquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Cal\`o, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, J\"org Tiedemann, Timothee Mickus
PlanSightRAG is a visual-first multimodal retrieval‑augmented generation system designed to automate question answering and compliance checking of civil standard plans. It processes plan imagery directly, using a ColNomic‑3B multi‑vector retrieval engine, an agentic Planner‑Retriever‑Auditor‑Synthesizer, and MaxSim heatmaps to provide an evidence trail. The system achieves high recall on a new 4,056‑pair benchmark from five state DOTs, and demonstrates near‑perfect verdict accuracy on synthetic compliance drawings when a rule threshold is supplied, outperforming OCR‑based baselines.
By Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty, Shivanand Venkanna Sheshappanavar
The paper introduces the Necessary Tool‑Evidence Path (NTEP) annotation scheme and its associated reward mechanism (NTEP‑R) to better supervise vision‑language models that use external tools. By explicitly specifying which evidence is needed and penalizing redundant tool calls, the authors train an 8B‑parameter model that shows improved accuracy and tool‑use efficiency across seven image‑grounded benchmarks. The approach demonstrates that fine‑grained supervision of tool‑evidence paths is essential for robust agentic VLM performance.
By Xingming Long, Yu Liu, Zhiwei Yang, Hanqi Feng, Shaojie Zhang, Barnabas Poczos, Chao Jiang, Zhenbo Luo, Lei Jiang, Pei Fu
arXiv:2602. 03045v2 Announce Type: replace Abstract: Large language models have recently enabled text-to-CAD systems that synthesize parametric CAD programs (e.
By Bo Yuan, Zelin Zhao, Petr Molodyk, Bin Hu, Yongxin Chen
arXiv:2607. 00089v1 Announce Type: new Abstract: Mechanistic interpretability has produced a rich inventory of component-level analyses that characterise what neural-network components encode and how they interact.
By Hussein Chouman, Wataru Sasaki, Tomokazu Matsui, Hirohiko Suwa, Keiichi Yasumoto
arXiv:2608. 15032v1 Announce Type: cross Abstract: Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state.
By Bruno Chicelli, Henrique Alves, Rodrigo Anselmo, Joshua Weinberg, Felipe Lemos, Jan Baryla
ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.
By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang