ExBind is a controlled diagnostic benchmark that isolates the visual‑to‑executable correspondence layer in multimodal coding and editing systems. It generates 250 broad and 240 targeted cases across SVG, DOM, canvas, tree, graph, and table formats, each with deterministic mappings to executable references. Models are evaluated solely on their ability to output the correct reference, with structural constraints scored without requiring reasoning traces.
By Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao, Jinli Suo
arXiv:2609.39380v1 Announce Type: new
Abstract: Reliable infographic generation requires facts, symbols, and visual relations to remain consistent through rendering and revision. Correcting one eleme...
By Yifan Li, Tong Li, Qi Zeng, Lishuai Gao, Ruwei Pan, Cong Wei, Shaohua Kevin Zhou, Zhuoliang Kang, Xiaoming Wei
arXiv:2609.08657v1 Announce Type: cross
Abstract: Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This str...
By Xiaochuan Zhong, Yifan Hou, Chenxi Pang, Shaobo Cui
arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.
By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv:2607. 19056v1 Announce Type: new Abstract: Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone.
By Yug Aditi Gupta, Prannay Hebbar
The paper introduces LayerWiseBench, a benchmark that evaluates visual language models on layer-wise chart understanding and editing. It focuses on three core concepts—layer attribution, layer binding, and visibility ordering—by pairing rendered charts with spatially aligned per-layer RGBA assets and functional role labels. The benchmark includes 2,800 charts, 7,329 understanding questions, and 53,791 editing variants, revealing that models excel at attribution and binding but struggle with visibility ordering, especially when editing overlapping components.
arXiv:2606. 00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle.
By Kai Xu, Ellis Brown, Shrikar Madhu, Rob Fergus, He He, Saining Xie
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
arXiv:2608. 16805v1 Announce Type: cross Abstract: Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance.
By Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Sixue Lin
MatToolBench is a new benchmark that evaluates multimodal GUI agents on professional materials science software. It contains 204 tasks across 10 tools in three modalities—GUI operation, OriginPro scripting, and code-based database queries—executed inside a Windows 11 VM. The benchmark offers fine-grained, expert-decomposed scoring and a high-performing multimodal judge for aesthetic assessment, revealing that strong general benchmark performance does not transfer to scientific workflows.
By Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li, Tianfan Fu, Bo Chen, Kai Yu, Xin Chen, Lu Chen
arXiv:2510.17932v5 Announce Type: replace-cross
Abstract: We introduce Chart2Code, a new benchmark for evaluating the chart understanding and code generation capabilities of large multimodal models (...
By Jiahao Tang, Henry Hengyuan Zhao, Lijian Wu, Zijian Zhang, Yifei Tao, Dongxing Mao, Yang Wan, Jingru Tan, Min Zeng, Min Li, Alex Jinpeng Wang
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class instance. Generic visual-question-answering accuracy marks the...