arXiv Computer Vision
Sep 2

ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence

ExBind is a controlled diagnostic benchmark that isolates the visual‑to‑executable correspondence layer in multimodal coding and editing systems. It generates 250 broad and 240 targeted cases across SVG, DOM, canvas, tree, graph, and table formats, each with deterministic mappings to executable references. Models are evaluated solely on their ability to output the correct reference, with structural constraints scored without requiring reasoning traces.

By Ziqian Wang, Yuxiao Cheng, Tingxiong Xiao, Jinli Suo
arXiv AI
Jul 1

An Executable Benchmarking Suite for Tool-Using Agents

arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.

By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
Hugging Face Trending Papers
Sep 8

Charts Are Beyond Pixels: Probing for Layer-Wise Chart Understanding and Editing

The paper introduces LayerWiseBench, a benchmark that evaluates visual language models on layer-wise chart understanding and editing. It focuses on three core concepts—layer attribution, layer binding, and visibility ordering—by pairing rendered charts with spatially aligned per-layer RGBA assets and functional role labels. The benchmark includes 2,800 charts, 7,329 understanding questions, and 53,791 editing variants, revealing that models excel at attribution and binding but struggle with visibility ordering, especially when editing overlapping components.