arXiv:2606. 31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents.
By Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig
PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.
By Xiaoqiu Wang, Yizhe Chi, Wenyi Li, Deyao Hong, Zhihan Shan, Mingju Gao, Kaisen Yang, Youjie Zheng, Calvin Xiao, Qinhuai Na
EditPPT is a multi‑agent framework that turns slide editing into a constrained tool‑selection task, using the native PowerPoint COM interface to perform localized shape‑level operations. By separating validation across modalities, its dual‑modal validators assess both instruction fidelity and visual quality, achieving high execution and accuracy rates even on long decks. The authors also introduce DeckEdit‑Bench, a benchmark of 28 human‑authored decks with 582 slides and 183 editing prompts across varying deck lengths.
By Jiheon Kim, Kyudan Jung, Jaegul Choo
ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.
By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
SLIDEFORGE is an LLM‑driven agent designed for controllable editing of slide decks while preserving layout, style, component structure, and native editability. It constructs a Deck State Graph that links visual decomposition, PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose a comprehensive evaluation framework measuring component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms direct prompting, screenshot‑based agents, and generic code‑agent baselines.
SLIDEFORGE is a new AI agent designed for controllable editing of presentation slides. It constructs a Deck State Graph that links visual decomposition, native PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose an evaluation framework that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms existing prompting, screenshot‑based, and generic code‑agent baselines.
By Haozhen Zheng, Fulin Wang, Tianhu Xiong, Yingjie Yu, Shengyi Qian, Hanchao Yu, Alex Schwing, Klara Nahrstedt, Mingyuan Wu