arXiv AI By Beatrice Alessandra Motetti, Emilien Guandalino, Daniele Jahier Pagliari, Alessio Burrello, Lorenz K. M\"uller, Konstantin Berestizshevsky, Lukas Cavigelli

Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

Read the original on arXiv AI →

arXiv:2608. 14446v1 Announce Type: new Abstract: In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

arXiv:2605. 29861v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have advanced autonomous agents from deep search, which retrieves concise factual answers, to deep research, which synthesizes scattered evidence into long-form reports.

By Chenghao Zhang, Guanting Dong, Yufan Liu, Tong Zhao, Xiaoxi Li, Zhicheng Dou
arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv Computation and Language
Sep 17

ReFigBench: Benchmarking Scientific Figure Reconstruction as Editable PowerPoint Artifacts

ReFigBench is a benchmark that evaluates how well multimodal coding agents can transform scientific overview figures into editable PowerPoint slides, preserving text, layout, and document structure. The study uses 1,000 real figures from arXiv, testing agents from four model families across two workflows—direct code generation and a specialized PPTX workflow—within ten different harness configurations. Evaluation combines deterministic artifact checks, automated scoring by judges, and blinded human comparisons, revealing that workflow and harness choices significantly affect reconstruction quality and that even the best agents fall short of the ideal rubric.

By Liyang Fan, Chi Wei, Yitai Li, Xinping Bi, Guhong Chen, Chenghao Sun, Haoxiang Yang, Qingwen Li, Kai Yan, Hong Li, Bo Li
arXiv Computation and Language
Aug 31

LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages

LandingAgent is a new framework for generating landing pages that are tailored to a specific target. It uses a reference‑annotated dataset called LandingBench, which abstracts real landing pages into structured elements such as section sequences, layout patterns, tone descriptors, visual emphasis, and CTA structure. The agentic framework operates in three phases—profiling the target, building a reference‑guided wireframe, and refining the page through critique—resulting in pages that are more faithful to the target, concise, readable, aesthetically pleasing, and structurally diverse compared to direct prompting.

By Injun Baek, HyeongSeok Lee, Yearim Kim, Junhoo Lee, Nojun Kwak