PPTArena: A Benchmark for PowerPoint Editing
arXiv:2512. 03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions.
PPTBench is a new benchmark that tests coding agents’ ability to reconstruct scientific flow diagrams from arXiv papers into editable PowerPoint slides. The dataset contains 500 tasks, each requiring agents to produce a single PPTX page with native, editable objects, and a four‑stage Agentic Judge evaluates validity, semantic correctness, rendering quality, and fine‑grained visual quality. Across 31 model configurations, the best score is 67.80, with a median of 19.47, showing that while agents can generate valid PPTX files, they still struggle with semantic and visual accuracy, especially text details.
arXiv:2512. 03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions.
SLIDEFORGE is a new AI agent designed for controllable editing of presentation slides. It constructs a Deck State Graph that links visual decomposition, native PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose an evaluation framework that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms existing prompting, screenshot‑based, and generic code‑agent baselines.
ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.
SLIDEFORGE is an LLM‑driven agent designed for controllable editing of slide decks while preserving layout, style, component structure, and native editability. It constructs a Deck State Graph that links visual decomposition, PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose a comprehensive evaluation framework measuring component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms direct prompting, screenshot‑based agents, and generic code‑agent baselines.
Code4Scene is a benchmark that evaluates coding agents on constructing and editing 3D scenes in Unreal Engine. It tests agents on two tasks: construction, where they must build a scene from open‑ended language, and editing, where they must recover a target scene from reference images while preserving everything else. The benchmark measures task fulfillment, artifact integrity, and physical validity, revealing that construction and editing performance are correlated but not interchangeable, with agents struggling most with spatial composition and precise edits.
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles.
arXiv:2605. 26144v2 Announce Type: replace-cross Abstract: We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents.
arXiv:2601.09487v2 Announce Type: replace Abstract: The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts...
arXiv:2609.36380v1 Announce Type: new Abstract: A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program...
arXiv:2605. 14398v3 Announce Type: replace Abstract: Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit physical constraints: contacts drift, shapes distort, and motion loses consistency.
Timeline-Bench is a benchmark comprising 56 real video‑editing tasks that require AI agents to transform raw production material into finished videos. Each task includes a brief, source assets, a container, and a set of tests that assess format, content, brief compliance, and quality based on 2,582 blind judgments by 43 video editors. In evaluations, the best agent resolved only 15 of the 56 tasks, and most failures were due to quality tests rather than technical errors.
arXiv:2608.24169v1 Announce Type: new Abstract: 3D geometry editing is a critical yet labor-intensive part of the graphics pipeline, requiring artists to translate creative intent into precise operat...