arXiv AI

ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback

arXiv AI
Aug 25

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
Hugging Face Trending Papers
Sep 2

SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

SLIDEFORGE is an LLM‑driven agent designed for controllable editing of slide decks while preserving layout, style, component structure, and native editability. It constructs a Deck State Graph that links visual decomposition, PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose a comprehensive evaluation framework measuring component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms direct prompting, screenshot‑based agents, and generic code‑agent baselines.

arXiv Computer Vision
Sep 4

SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

SLIDEFORGE is a new AI agent designed for controllable editing of presentation slides. It constructs a Deck State Graph that links visual decomposition, native PowerPoint object structure, and perceptual organization, enabling theme‑preserving reconstruction through slide‑native operations and rendered‑state verification. The authors also propose an evaluation framework that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability, and demonstrate that SLIDEFORGE outperforms existing prompting, screenshot‑based, and generic code‑agent baselines.

By Haozhen Zheng, Fulin Wang, Tianhu Xiong, Yingjie Yu, Shengyi Qian, Hanchao Yu, Alex Schwing, Klara Nahrstedt, Mingyuan Wu
arXiv Computation and Language
Sep 1

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.

By Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
arXiv AI
Sep 2

ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation

ReDeck introduces a step‑level render‑grounded refinement framework for document‑to‑slide generation, breaking slide revision into atomic edit actions with immediate renderer‑derived observations. It employs multi‑granular feedback—step‑level spatial checks, turn‑level adaptive critique, and a submission‑level layout gate—to balance local repair with overall quality. The authors also present DeckQuiz, a benchmark that separates content fidelity, spatial correctness, and design quality, and demonstrate ReDeck’s superior performance across GPT‑5.4, Claude‑4.6, and Gemini‑3.1.

By Muzhao Tian, Zezi Zeng, Yifan Yang, Xin Gao, Yan Li, Zisu Huang, Xiaohua Wang, Changze Lv, Mingxi Cheng, Bei Liu, Kai Qiu, Qi Dai, Dong Chen, Yue Dong, Xiaoqing Zheng, Ji Li, Chong Luo
arXiv Machine Learning
Jun 11

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

arXiv:2601. 04203v2 Announce Type: replace-cross Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generation with multi-modal feedback.

By Xueqing Wu, Zihan Xue, Da Yin, Shuyan Zhou, Kai-Wei Chang, Nanyun Peng, Yeming Wen
Hugging Face Trending Papers
Aug 2

TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents

Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time.

arXiv AI
Jul 7

Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning

arXiv:2603. 02070v3 Announce Type: replace Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facilitate an iterative reasoning and elicitation process, where the human's role is to guide the AI planner according to their preferences and expertise.

By Guilhem Fouilh\'e, Rebecca Eifler, Antonin Poch\'e, Sylvie Thi\'ebaux, Nicholas Asher