arXiv AI By Haodong Chen, Xuanhe Zhou, Wei Zhou, Xinyue Shao, Yanbing Zhu, Bo Wang, Jiawei Hong, Anya Jia, Fan Wu

X+Slides: Benchmarking Audience-Conditioned Slide Generation

Read the original on arXiv AI →

arXiv:2606. 19256v1 Announce Type: new Abstract: Automatically generating slide decks from source documents is an important application of large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

SlideLab: Audience-Centered Scientific Slide Generation and Evaluation

SlideLab is a training‑free, multi‑agent framework that generates scientific presentations directly from research papers. It first plans a coherent narrative, then iteratively builds and refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab outperformed both open‑source and commercial systems on 77% of papers while using about four times fewer inference tokens than the strongest open‑source baseline. The authors also introduce ConfArena, an audience‑oriented evaluation framework that simulates a conference room and assesses presentations slide by slide, matching human system rankings and detecting issues such as falsified numbers, degraded figures, dropped slides, and shuffled slide order.

By Vidushee Vats, Karun Sharma, Yuxia Wang
arXiv AI
Jul 1

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

arXiv:2606. 31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents.

By Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig
arXiv AI
Aug 25

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
arXiv Computation and Language
Sep 25

What, When, and How: Audio Description as Constrained Global Optimization

The paper presents a new approach to automatic audio description (AD) that treats the task as a constrained global optimization problem. It jointly decides what visual content is narratively important, when it can be spoken without overlapping dialogue, and how to phrase it within time limits. Using large language models for salience estimation and a mixed‑integer linear program for scheduling, the system outperforms prior methods on the REFRAMED benchmark, especially in temporal placement and narrative relevance.

By Igor Sterner, Mirella Lapata, Alex Lascarides, Frank Keller
arXiv Machine Learning
Sep 11

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

The paper introduces REVA, a method for compressing retrieval-augmented generation (RAG) prompts by aggregating historical query–document–model interactions into reusable evidence views. REVA mines attention traces from the target generator, maps token-level attention to readable words, aggregates importance across repeated document accesses, and produces budget‑specific plain‑text views that maintain document order and the standard RAG interface. Experiments on four benchmarks with modern LLMs show that REVA improves generation quality by 1.0–5.8 points over existing compressors while reducing compression overhead by 5.3 to 15.6 times and adding less than 40 ms of latency.

By Tuan Nguyen, Qiran Hu, Banruo Liu, Khoa D. Doan, Kok-Seng Wong, Fan Lai