ExBind: A Controlled Diagnostic Benchmark for Visual-to-Executable Correspondence
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
ExBind is a controlled diagnostic benchmark that isolates the visual‑to‑executable correspondence layer in multimodal coding and editing systems. It generates 250 broad and 240 targeted cases across SVG, DOM, canvas, tree, graph, and table formats, each with deterministic mappings to executable references. Models are evaluated solely on their ability to output the correct reference, with structural constraints scored without requiring reasoning traces.
arXiv:2609.39380v1 Announce Type: new Abstract: Reliable infographic generation requires facts, symbols, and visual relations to remain consistent through rendering and revision. Correcting one eleme...
arXiv:2609.08657v1 Announce Type: cross Abstract: Charts are structured visual compositions whose elements have distinct functional roles, semantic correspondences, and visibility relations. This str...
arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.
arXiv:2607. 19056v1 Announce Type: new Abstract: Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone.
The paper introduces LayerWiseBench, a benchmark that evaluates visual language models on layer-wise chart understanding and editing. It focuses on three core concepts—layer attribution, layer binding, and visibility ordering—by pairing rendered charts with spatially aligned per-layer RGBA assets and functional role labels. The benchmark includes 2,800 charts, 7,329 understanding questions, and 53,791 editing variants, revealing that models excel at attribution and binding but struggle with visibility ordering, especially when editing overlapping components.