Multimodal coding and editing systems must map a visible or semantic referent to the exact executable object that can be edited. A wrong reference may select a valid but incorrect DOM node, SVG elemen...
arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.
By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv:2609.08657v1 Announce Type: cross
Abstract: Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This str...
By Xiaochuan Zhong, Yifan Hou, Chenxi Pang, Shaobo Cui
arXiv:2609.39380v1 Announce Type: new
Abstract: Reliable infographic generation requires facts, symbols, and visual relations to remain consistent through rendering and revision. Correcting one eleme...
By Yifan Li, Tong Li, Qi Zeng, Lishuai Gao, Ruwei Pan, Cong Wei, Shaohua Kevin Zhou, Zhuoliang Kang, Xiaoming Wei
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
arXiv:2607. 19056v1 Announce Type: new Abstract: Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone.
By Yug Aditi Gupta, Prannay Hebbar
arXiv:2606. 00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe what it sees and name the underlying pattern, yet still fail to choose the matching candidate.
By Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng, Qiyao Sun, Xuanyu Ji, Qingyong Hu
The paper introduces LayerWiseBench, a benchmark that evaluates visual language models on layer-wise chart understanding and editing. It focuses on three core concepts—layer attribution, layer binding, and visibility ordering—by pairing rendered charts with spatially aligned per-layer RGBA assets and functional role labels. The benchmark includes 2,800 charts, 7,329 understanding questions, and 53,791 editing variants, revealing that models excel at attribution and binding but struggle with visibility ordering, especially when editing overlapping components.
The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.
By Guangzhe Zhang
arXiv:2607. 18316v1 Announce Type: cross Abstract: Tool-augmented language-model agents execute multi-step workflows over external systems, resolving an entity once and then acting on it across subsequent steps.
By Rahul Suresh Babu, Shashank Indukuri
arXiv:2606. 08151v1 Announce Type: new Abstract: Tool-using LLM agents often fail not because relevant text is absent, but because decisive evidence is not selected, compressed, or surfaced at action time.
By Xinyu Guan, Qianyang Zhao, Yuming Deng
arXiv:2605.31351v2 Announce Type: replace-cross
Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradig...
By Yi Zhao, Siqi Wang, Zhe Hu, Yushi Li, Jing Li