SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.
By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
SLIDEFORGE is an LLM‑driven agent designed for controllable editing of slide decks while preserving layout, style, component structure, and native editability. It constructs a Deck State Graph that links visual decomposition, PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose a comprehensive evaluation framework measuring component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms direct prompting, screenshot‑based agents, and generic code‑agent baselines.
arXiv:2604. 19971v2 Announce Type: replace-cross Abstract: Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking.
By Xuxin Tang, Ibrahim Tahmid, Eric Krokos, Kirsten Whitley, Xuan Wang, Chris North
SLIDEFORGE is a new AI agent designed for controllable editing of presentation slides. It constructs a Deck State Graph that links visual decomposition, native PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose an evaluation framework that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms existing prompting, screenshot‑based, and generic code‑agent baselines.
By Haozhen Zheng, Fulin Wang, Tianhu Xiong, Yingjie Yu, Shengyi Qian, Hanchao Yu, Alex Schwing, Klara Nahrstedt, Mingyuan Wu
PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.
By Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
ReDeck introduces a step‑level render‑grounded refinement framework for document‑to‑slide generation, breaking slide revision into atomic edit actions with immediate renderer‑derived observations. It employs multi‑granular feedback—step‑level spatial checks, turn‑level adaptive critique, and a submission‑level layout gate—to balance local repair with overall quality. The authors also present DeckQuiz, a benchmark that separates content fidelity, spatial correctness, and design quality, and demonstrate ReDeck’s superior performance across GPT‑5.4, Claude‑4.6, and Gemini‑3.1.
By Muzhao Tian, Zezi Zeng, Yifan Yang, Xin Gao, Yan Li, Zisu Huang, Xiaohua Wang, Changze Lv, Mingxi Cheng, Bei Liu, Kai Qiu, Qi Dai, Dong Chen, Yue Dong, Xiaoqing Zheng, Ji Li, Chong Luo
arXiv:2601.12208v2 Announce Type: replace
Abstract: Evaluating conversational systems in multi-turn settings remains a fundamental challenge. Conventional pipelines typically rely on manually defined...
By Yunzhe Li, Richie Yueqi Feng, Tianxin Wei, Chin-Chia Hsu
arXiv:2601. 04203v2 Announce Type: replace-cross Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generation with multi-modal feedback.
By Xueqing Wu, Zihan Xue, Da Yin, Shuyan Zhou, Kai-Wei Chang, Nanyun Peng, Yeming Wen
Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time.
arXiv:2608.29621v1 Announce Type: cross
Abstract: Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt con...
By Junxiang Liu, Lin Wang, Haiyu Shi, Hongxu Ma, Xiaoyu Yang, Chunjie Chen, Xiaoxiao Xu, Kaiqiao Zhan, Boao Wang, Shuizhou Shi, Tianyun Zhu, Jie Li, Jiangtong Li
arXiv:2603. 02070v3 Announce Type: replace Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise.
By Guilhem Fouilh\'e, Rebecca Eifler, Antonin Poch\'e, Sylvie Thi\'ebaux, Nicholas Asher
arXiv:2609.14677v1 Announce Type: new
Abstract: Large language models (LLMs) have changed the way people engage with stories. Drawing on public chatbot logs, we can see that when users generate stori...
By Advait Deshmukh, Nora Benedict, Melanie Walsh, Maria Antoniak