arXiv Computer Vision

VectorHarness: Recovering Editable, Relation-Preserving Structure from Scientific Graphics

arXiv AI
Sep 2

Figures as Programs: Recursive Generation of Editable Scientific Figures

The paper introduces “FigTree”, a multi-agent system that automatically converts a scientific paper into a structured vector figure by recursively constructing SVG programs. It decomposes figures into hierarchical regions, generates each region as a short SVG program, and assembles them, using a render‑critic refinement loop to trace and repair visual defects. Evaluations show that “FigTree” produces high‑quality figures and allows more effective editing than raster‑based methods.

By Yepeng Liu, Dasen Dai, Chengzhi Liu, Yiren Song, Hai Ci, Yu Zhang, Qi Zhang, Mike Zheng Shou, Xin Eric Wang, Yuheng Bu
arXiv Machine Learning
Aug 28

Chart2SVG: Editable SVG Generation from Raster Chart Images

Chart2SVG is a multimodal large language model that transforms static raster chart images into editable SVGs enriched with semantic structure. By embedding chart‑specific semantic tokens into a vision‑language framework and training on the Beagle+ dataset of 33K distilled chart samples, the model captures both geometric primitives and their functional roles. The resulting SVGs are visually accurate and structurally consistent, and the accompanying Chart Structure Graph (CSG) exposes visual dependencies for interactive exploration, chart repurposing, and layout reuse.

By Jinning Cui, Lu Chen, Haoyan Shi, Yue He, Chenglong Wang, Mengyu Zhou, Weidong Huang, Yunhai Wang
arXiv Computation and Language
3d ago

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.

By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
arXiv Computer Vision
Sep 4

SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

SLIDEFORGE is a new AI agent designed for controllable editing of presentation slides. It constructs a Deck State Graph that links visual decomposition, native PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose an evaluation framework that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms existing prompting, screenshot‑based, and generic code‑agent baselines.

By Haozhen Zheng, Fulin Wang, Tianhu Xiong, Yingjie Yu, Shengyi Qian, Hanchao Yu, Alex Schwing, Klara Nahrstedt, Mingyuan Wu
Hugging Face Trending Papers
Sep 2

SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

SLIDEFORGE is an LLM‑driven agent designed for controllable editing of slide decks while preserving layout, style, component structure, and native editability. It constructs a Deck State Graph that links visual decomposition, PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose a comprehensive evaluation framework measuring component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms direct prompting, screenshot‑based agents, and generic code‑agent baselines.

arXiv Computer Vision
Sep 1

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

arXiv:2607.02290v2 Announce Type: replace Abstract: Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a kno...

By Zhaokai Wang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Yiguo He, Mohan Zhang, Leyao Gu, Yan Li, Xiangyu Zhao, Ning Liao, Shaofeng Zhang, Xuanhe Zhou, Zhihang Zhong, Xue Yang