The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi‑layer document and UI designs. It presents CoDeLayout, a VQA dataset of about 20,000 real‑world layouts annotated with compositional element pairs and design intent. The authors identify semantic drift and structural ambiguity as key challenges for vision‑language models and propose MASON, a post‑training approach that combines multimodal alignment and structural perception to improve performance, achieving 91.66% accuracy with only 30% of the training data.
By Yiyang Huang, Zhaowen Wang, Simon Jenni, Jing Shi, Yitian Zhang, Yizhou Wang, Yun Fu
The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi-layer document and UI layouts that involve hierarchical relationships among visually entangled elements. It presents CoDeLayout, a VQA dataset of about 20,000 real-world layouts annotated with compositional element pairs and design intent, and identifies two main challenges for current vision‑language models: semantic drift between textual metadata and visual content, and structural ambiguity in hierarchical inter‑element relationships. To address these, the authors propose MASON, a post‑training paradigm that combines multimodal alignment and structural perception, achieving a 91.66% accuracy on CoDeLayout and outperforming full‑data direct fine‑tuning with only 30% of the training data.
The paper introduces VLM-CAD, a workflow that uses Vision Language Models (VLMs) for analog circuit sizing while mitigating spatial blindness and logical hallucinations. It incorporates a neuro‑symbolic parsing module, Image2Net, to convert schematics into topological graphs and JSON, and an Explainable Trust Region Bayesian Optimization method, ExTuRBO, to guide design decisions with sensitivity evidence. Experiments on 12 sizing tasks across six circuits and four technology platforms show a Strict Pass@1 of 23.3% and a Relaxed Pass@1 of 91.7%.
By Guanyuan Pan, Shuai Wang, Yugui Lin, Tiansheng Zhou, Pietro Li\`o, Zhenxin Zhao, Yaqi Wang
arXiv:2608. 04726v1 Announce Type: new Abstract: Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels.
By Yongxin Wang, Ruizhe Zhou, Yueling Tang, Yingying Zhu, Xuemin Zhao, Xiaojun Chang, Xiaodan Liang
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
arXiv:2607. 15418v1 Announce Type: new Abstract: We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings -- a core media in architecture, civil, and many other engineering practices.
By Yoonhwa Jung, Junryu Fu, Mani Golparvar-Fard
arXiv:2606. 05058v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models.
By Jingyuan Chen, Sheng Jin, Haopeng Sun, Wentao Liu, Chen Qian
arXiv:2605. 30794v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks.
By Qian Kou, Xiaofeng Shi, Yulin Li, Xiaosong Qiu, Xinyang Wang, Hua Zhou, Cao Dongxing
arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.
By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
UReason is a benchmark that evaluates how well unified multimodal models (UMMs) align textual reasoning with image generation. It contains 2,000 human‑curated instances across five reasoning‑intensive tasks—Code, Arithmetic, Spatial, Attribute, and Text—and compares direct generation, reasoning‑guided generation, and decontextualized generation. The study finds that while reasoning‑guided generation improves over direct generation, decontextualized generation consistently outperforms it, indicating that the visual semantics in textual reasoning are not reliably reflected in the generated images.
By Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick
arXiv:2608. 09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image.
By Xinqi Yang, Kang An, Tengyue Wang, Zhongyu Yang, Chenxu Du, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren
arXiv:2606. 10833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) demonstrate strong performance on general multimodal reasoning benchmarks, yet their ability to perform engineering reasoning remains largely unexplored.
By Syed Wasiq, Syed Mohamad Tawseeq, Yashwant Pravinrao Bangde, Debaditya Roy