arXiv Machine Learning By Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni

THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout

Read the original on arXiv Machine Learning →

THEIA is a multimodal dataset that pairs thousands of analog circuit layout images with question‑answer conversations, and it introduces a benchmark using a fine‑tuned vision‑language model to analyze GDSII files. The dataset and benchmark enable designers to interact with and query physical layouts as intuitive, meaningful entities. Experiments on five realistic tasks show the fine‑tuned model outperforms general‑purpose vision‑language models by up to 73%, revealing a significant gap between general multimodal reasoning and domain‑specific layout understanding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 28

Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models

The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi‑layer document and UI designs. It presents CoDeLayout, a VQA dataset of about 20,000 real‑world layouts annotated with compositional element pairs and design intent. The authors identify semantic drift and structural ambiguity as key challenges for vision‑language models and propose MASON, a post‑training approach that combines multimodal alignment and structural perception to improve performance, achieving 91.66% accuracy with only 30% of the training data.

By Yiyang Huang, Zhaowen Wang, Simon Jenni, Jing Shi, Yitian Zhang, Yizhou Wang, Yun Fu
Hugging Face Trending Papers
Aug 27

Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models

The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi-layer document and UI layouts that involve hierarchical relationships among visually entangled elements. It presents CoDeLayout, a VQA dataset of about 20,000 real-world layouts annotated with compositional element pairs and design intent, and identifies two main challenges for current vision‑language models: semantic drift between textual metadata and visual content, and structural ambiguity in hierarchical inter‑element relationships. To address these, the authors propose MASON, a post‑training paradigm that combines multimodal alignment and structural perception, achieving a 91.66% accuracy on CoDeLayout and outperforming full‑data direct fine‑tuning with only 30% of the training data.

arXiv AI
Sep 18

VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizing

The paper introduces VLM-CAD, a workflow that uses Vision Language Models (VLMs) for analog circuit sizing while mitigating spatial blindness and logical hallucinations. It incorporates a neuro‑symbolic parsing module, Image2Net, to convert schematics into topological graphs and JSON, and an Explainable Trust Region Bayesian Optimization method, ExTuRBO, to guide design decisions with sensitivity evidence. Experiments on 12 sizing tasks across six circuits and four technology platforms show a Strict Pass@1 of 23.3% and a Relaxed Pass@1 of 91.7%.

By Guanyuan Pan, Shuai Wang, Yugui Lin, Tiansheng Zhou, Pietro Li\`o, Zhenxin Zhao, Yaqi Wang