arXiv Computation and Language

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.

arXiv AI
Aug 25

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
arXiv Computation and Language
3d ago

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.

By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
arXiv AI
5d ago

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

VisInteract introduces a new paradigm for Text-to-Visualization that treats imperfect user queries as a dynamic, interaction-driven problem, requiring the system to recover true intent through multi-turn dialogue. The authors present VisInteract-Bench, the first benchmark for interactive Text-to-Vis, featuring controlled imperfection injection, a realistic user agent, and dual-perspective automated evaluation. They also propose Vis-MCTS, an enhanced Monte Carlo Tree Search algorithm that incorporates progressive widening, cross-rollout information sharing, and dimension-aware reward decomposition, achieving significant performance gains over existing baselines.

By Wenxin Xu, Jinwei Lu, Hwanhee Kim, Chen Jason Zhang, Xiao-Yong Wei, Haoyang Li, Yuanfeng Song
arXiv AI
Jun 9

IMUG-Bench: Benchmarking Unified Multimodal Models on Interleaved Understanding and Generation

arXiv:2606. 09169v1 Announce Type: new Abstract: In recent years, unified multimodal models (UMMs) have emerged to support both understanding and generation within a single framework.

By Lingyi Meng, Zecong Tang, Haoran Li, Tengju Ru, Zhejun Cui, Weitong Lian, Qi Kang, Hangshuo Cao, Yichen Zhu, Yechi Liu, Kaixuan Wang, Yu-Jie Yuan, Chunwei Wang, Yu Zhang, Bo Dai
arXiv Machine Learning
Jun 11

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

arXiv:2601. 04203v2 Announce Type: replace-cross Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generation with multi-modal feedback.

By Xueqing Wu, Zihan Xue, Da Yin, Shuyan Zhou, Kai-Wei Chang, Nanyun Peng, Yeming Wen