arXiv AI

BlueprintAgent: Constraint-Triggered Targeted Revisits for Simulation-Ready Generation from Scanned Structural Blueprints

BlueprintAgent (BPA) is a multimodal agent that converts scanned reinforced‑concrete building blueprints into simulation‑ready frame models by treating a large language model (MLLM) as the primary reader and decision maker, while OCR and computer vision provide localized evidence. BPA implements engineering constraints as callable validators that trigger targeted MLLM revisits over specific regions, enabling precise beam‑column support, span count, and 3D continuity checks. In evaluations on 300 real scanned sheets from 20 projects, BPA achieved a macro‑averaged Beam F1 of 0.994, vastly outperforming single‑MLLM zero‑shot (0.301) and fixed‑pipeline (0.820) baselines.

arXiv AI
Aug 19

Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement

The paper introduces a novel framework that converts structural framing plans from PDF drawings into editable finite‑element model drafts without requiring task‑specific detector training. It combines a deterministic geometry extraction stage—estimating scale, recognizing five entity classes, and assembling a drafting grammar—with an agentic vision‑language refinement stage that proposes corrections, performs admission tests, and ensures fail‑closed transactions. Evaluation on a 100‑plan benchmark shows high accuracy, with scale within 0.1% and recall/precision values above 0.86 for all component types.

By Mohammad Talebi-Kalaleh, Qipei Mei
arXiv AI
Jul 17

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

arXiv:2607. 14896v1 Announce Type: cross Abstract: Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report.

By Sizhong Qin, Yi Gu, Yao Jiang, Ao Cai, Changjian Zhou, Shaoxuan Shuai, Jiachang Wang, Tianhao Shen, Yueqiang Li, Xinhao Li, Li Zeng, Yueshi Chen, Dachen Gao, Genrong Xu, Wenjie Liao, Xinzheng Lu
arXiv Computer Vision
Aug 27

Training-Free Agentic Computer Vision for Structural Component Detection in 2D Structural Framing Plans

The paper presents a training‑free, agentic computer‑vision system that converts 2D structural framing plan PDFs into editable finite‑element model drafts. It uses a deterministic stage to extract geometric primitives, estimate scale, and recognize five entity classes via a drafting grammar, followed by an agentic stage that applies typed corrections and fail‑closed transactions. Evaluation on a 100‑plan benchmark shows high precision and recall across columns, beams, walls, braces, and openings, with scale estimates within 0.1% of reference.

By Mohammad Talebi-Kalaleh, Qipei Mei
arXiv AI
Sep 16

From Transient Prompts to Persistent Control: Scientific Poster Generation via Recursive Semantic-Geometric Contracts

The paper introduces PosterVisor, a framework that replaces transient prompts with persistent control for scientific poster generation. It uses an Orchestrator to compile rubrics into Semantic‑Geometric Contracts that bind claims, sources, visuals, budgets, and spatial commitments, and employs Recursive Contract Enforcement to dynamically trigger checks and prevent silent regressions. Implementations in HTML/CSS and PPTX show improved QA accuracy and higher human preference on benchmark datasets.

By Runze Li, Yukun Zhao, Can Xu, Yucheng Shen, Shuaiqiang Wang, Jianmin Wu, Lingyong Yan, Dawei Yin
arXiv AI
Aug 11

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

arXiv:2608. 09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion.

By Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren
arXiv AI
Jun 17

Blueprint First, Model Second: A Framework for Deterministic LLM Workflow

arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.

By Libin Qiu, Yuhang Ye, Zhirong Gao, Xide Zou, Junfu Chen, Ziming Gui, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Kun Zhao
Hugging Face Trending Papers
Jun 24

Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints

Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficiently understood. This paper reports a reproducible phenomenon observed in a production Agent system: when Tool Calling and JSON Schema constraints are simultaneously enabled, multiple open-weight models cease invoking tools despite maintaining high schema compliance.

arXiv Computation and Language
Aug 27

PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

PlanSightRAG is a visual-first multimodal retrieval‑augmented generation system designed to automate question answering and compliance checking of civil standard plans. It processes plan imagery directly, using a ColNomic‑3B multi‑vector retrieval engine, an agentic Planner‑Retriever‑Auditor‑Synthesizer, and MaxSim heatmaps to provide an evidence trail. The system achieves high recall on a new 4,056‑pair benchmark from five state DOTs, and demonstrates near‑perfect verdict accuracy on synthetic compliance drawings when a rule threshold is supplied, outperforming OCR‑based baselines.

By Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty, Shivanand Venkanna Sheshappanavar
arXiv AI
Aug 18

Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints

arXiv:2608. 15032v1 Announce Type: cross Abstract: Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, dimensional and multi-hop reasoning, and grounding in construction conventions that the drawings never state.

By Bruno Chicelli, Henrique Alves, Rodrigo Anselmo, Joshua Weinberg, Felipe Lemos, Jan Baryla