arXiv Computation and Language By Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

Read the original on arXiv Computation and Language →

PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 25

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

SlideGen is a collaborative vision‑language multi‑agent framework designed to generate scientific presentation slides from research papers. It assigns specialized agents to outline the presentation structure, align figures and tables with key claims, generate speaker notes, and compose editable PPTX slides using a diverse layout library. The system introduces a geometry‑aware density metric to evaluate visual clutter and demonstrates significant improvements in layout balance, content coverage, and text coherence over existing baselines on a 200‑paper benchmark.

By Xin Liang, Zhilin Zhang, Xiang Zhang, Haoran Su, Yiwei Xu, Siqi Sun, Chenyu You
arXiv Computation and Language
3d ago

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.

By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
arXiv AI
5d ago

VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

VisInteract introduces a new paradigm for Text-to-Visualization that treats imperfect user queries as a dynamic, interaction-driven problem, requiring the system to recover true intent through multi-turn dialogue. The authors present VisInteract-Bench, the first benchmark for interactive Text-to-Vis, featuring controlled imperfection injection, a realistic user agent, and dual-perspective automated evaluation. They also propose Vis-MCTS, an enhanced Monte Carlo Tree Search algorithm that incorporates progressive widening, cross-rollout information sharing, and dimension-aware reward decomposition, achieving significant performance gains over existing baselines.

By Wenxin Xu, Jinwei Lu, Hwanhee Kim, Chen Jason Zhang, Xiao-Yong Wei, Haoyang Li, Yuanfeng Song