arXiv AI

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv AI
6d ago

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

The paper surveys multimodal speculative decoding, examining whether diffusion-based block‑parallel generative drafting—successful in text‑only LLMs—can be applied to Vision‑Language, Video‑Language, Audio, and Vision‑Language‑Action models. It introduces a taxonomy separating drafter‑side parallelism from other design choices, and presents an empirical comparison across benchmarks such as OCR, VQA, visual reasoning, and image captioning. The study highlights current limitations, outlines open challenges, and suggests future research directions for multimodal speculative decoding.

By Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan, Qiang Hui, Shuting Liu, Fuyuan Shi, Ting Lu, Shaoan Zhao, Xueqiang Guo, Xinpei Su, Jianbing Zhang, Xinyu Dai, Kai Wang, Shiguo Lian
arXiv AI
Jun 4

MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A

arXiv:2606. 04231v1 Announce Type: cross Abstract: Recent advances in multimodal retrieval-augmented generation (MM-RAG) have shifted toward minimal parsing, relying on page-level images for producing retriever embeddings and for answer generation.

By Hanoz Bhathena, Parin Rajesh Jhaveri, Rohan Mittal, Prateek Singh, Aymen Kallala, Rachneet Kaur, Yiqiao Jin, Zhen Zeng, Adwait Ratnaparkhi, Denis Kochedykov
arXiv AI
Jul 1

PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

arXiv:2606. 31154v1 Announce Type: cross Abstract: Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents.

By Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, Graham Neubig