arXiv:2601.09487v2 Announce Type: replace
Abstract: The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts...
By Yunqiao Yang, Wenbo Li, Houxing Ren, Zimu Lu, Ke Wang, Zhiyuan Huang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li
SlideLab is a training‑free, multi‑agent framework that generates scientific presentations directly from research papers. It first plans a coherent narrative, then iteratively builds and refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab outperformed both open‑source and commercial systems on 77% of papers while using about four times fewer inference tokens than the strongest open‑source baseline. The authors also introduce ConfArena, an audience‑oriented evaluation framework that simulates a conference room and assesses presentations slide by slide, matching human system rankings and detecting issues such as falsified numbers, degraded figures, dropped slides, and shuffled slide order.
By Vidushee Vats, Karun Sharma, Yuxia Wang
arXiv:2606. 31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents.
By Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig
SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.
By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.
By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
The paper introduces REVA, a method for compressing retrieval-augmented generation (RAG) prompts by aggregating historical query–document–model interactions into reusable evidence views. REVA mines attention traces from the target generator, maps token-level attention to readable words, aggregates importance across repeated document accesses, and produces budget‑specific plain‑text views that maintain document order and the standard RAG interface. Experiments on four benchmarks with modern LLMs show that REVA improves generation quality by 1.0–5.8 points over existing compressors while reducing compression overhead by 5.3 to 15.6 times and adding less than 40 ms of latency.
By Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai
arXiv:2512. 03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions.
By Michael Ofengenden, Yunze Man, Ziqi Pang, Liang-Yan Gui, Yu-Xiong Wang
WanPE is a 397‑B parameter prompt‑enhancement model that learns director‑level cinematic planning from 1.05 M real‑world videos. It generates shot‑level cinematic plans through video‑grounded reverse construction and uses Semantic‑Consistency GRPO (SC‑GRPO) to maintain user intent across shots and time. In evaluations, WanPE improves human preference over raw prompts by up to 50.86 points for 30‑second videos and outperforms commercial offerings for shorter durations.
By Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
The Spoken Wikipedia Presentation Corpus extends the existing Spoken Wikipedia Corpora by adding LLM-generated slide decks for multimodal automatic speech recognition (ASR). Slides are produced through a hybrid pipeline that combines LLM-based content planning with rule-based design, generating titles, bullet points, takeaway messages, and visual descriptions for illustrations. The corpus is evaluated with multiple ASR and spoken language models, achieving a best micro-WER of 10.23% and micro-CER of 6.48% on audio-only inputs, with English performing best and lower-resource languages showing higher error rates.
By Thomas Ranzenberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer
arXiv:2609.38406v1 Announce Type: new
Abstract: Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct mi...
By Eftekhar Hossain, John Salvador, Santu Karmaker
arXiv:2609.00226v1 Announce Type: new
Abstract: Automatic academic paper-to-slide generation is inherently iterative, because creating an effective presentation requires repeated cycles of generation...
By Tarik Can Ozden, Sachidanand VS, Furkan Horoz, Ozgur Kara, Dilek Hakkani-T\"ur, Junho Kim, James Matthew Rehg
arXiv:2606. 01252v1 Announce Type: cross Abstract: Multi-target cross-lingual text summarization (MTXLS), which summarizes a source document into multiple target languages, is increasingly important as users consume content in diverse languages, but remains underexplored.
By Sangwon Ryu, Yihong Liu, Mingyang Wang, Yunsu Kim, Jungseul Ok, Gary Geunbae Lee, Hinrich Schuetze