arXiv AI

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

arXiv:2608. 14446v1 Announce Type: new Abstract: In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with.

arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv Computation and Language
Sep 17

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.

By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
arXiv Computation and Language
Aug 31

LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

LandingAgent is a new framework for generating landing pages that are tailored to a specific target. It uses a reference‑annotated dataset called LandingBench, which abstracts real landing pages into structured elements such as section sequences, layout patterns, tone descriptors, visual emphasis, and CTA structure. The agentic framework operates in three phases—profiling the target, building a reference‑guided wireframe, and refining the page through critique—resulting in pages that are more faithful to the target, concise, readable, aesthetically pleasing, and structurally diverse compared to direct prompting.

By Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak
arXiv Computation and Language
Sep 18

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

The paper introduces Data Journalist Agent (Data2Story), a multi‑agent framework that orchestrates specialized roles into a single virtual newsroom to produce evidence‑grounded, multimodal news stories. It ensures every claim is traceable to data, code, or external references via an Inspector, and generates interactive visualizations such as maps and audio to match reader interests. Evaluations on 18 articles show competitive performance in angle coverage, rubric scores, and verifiability, while human writers still lead in editorial angle and creative design.

By Kevin Qinghong Lin, Batu EI, Yuhong Shi, Pan Lu, Juil Sock, Djordje Padejski, Philip Torr, James Zou
arXiv AI
Jul 8

Prompt-to-Paper: Agentic AI System for Bioinformatics

arXiv:2607. 05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication.

By Ramsha Kamran, Maheera Amjad, Zartasha Mustansar, Arsalan Shaukat, Salma Sherbaz, Muhammad U. S. Khan
arXiv AI
Sep 3

Multimodal Language Models as Text-to-Image Model Evaluators

Multimodal Language Models as Text-to-Image Model Evaluators presents MT2IE, a framework where a multimodal large language model generates evaluation prompts and scores images, achieving higher correlation with human judgment than prior metrics. MT2IE recovers official T2I model rankings using only 20 prompts—far fewer than traditional benchmarks—and adapts prompts to each model’s performance, maintaining informative scoring ranges. The approach demonstrates that dynamic, interactive evaluation can replace static benchmarks as T2I models improve.

By Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat, Koustuv Sinha, Melissa Hall, Amy Zhang, Michal Drozdzal, Adriana Romero-Soriano
arXiv AI
Aug 25

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP‑Bench proposes a new benchmark for evaluating agentic tool planning in multimodal large language models (MLLMs) that generate interleaved text-and-image responses. The benchmark contains 7,702 QA pairs, including 1,592 visual‑question‑answer pairs, across eight categories and 25 visual‑critical intents, all verified by humans. A Multi‑Agent MLLM‑as‑a‑Judge (MAM) system is introduced to assess tool‑call precision, missed opportunities, and overall response quality without relying on ground‑truth references.

By Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang