Multimedia and Visual Analytics in the Agentic Era
arXiv:2504. 06138v3 Announce Type: replace-cross Abstract: Professional users need tools to help them gain actionable insights from large multimedia collections.
arXiv:2504. 06138v3 Announce Type: replace-cross Abstract: Professional users need tools to help them gain actionable insights from large multimedia collections.
The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.
arXiv:2608.28637v1 Announce Type: new Abstract: Autonomous scientific discovery systems can generate large numbers of research ideas, experiments, and manuscripts with minimal human intervention. As...
New experimental AI tool helps people explore the context and origin of images seen online.
SlideLab is a training‑free, multi‑agent framework that generates scientific presentations directly from research papers. It first plans a coherent narrative, then iteratively builds and refines a shared slide deck using agents for content planning, visual generation, layout refinement, and grounding verification. In a blind human preference study, SlideLab outperformed both open‑source and commercial systems on 77% of papers while using about four times fewer inference tokens than the strongest open‑source baseline. The authors also introduce ConfArena, an audience‑oriented evaluation framework that simulates a conference room and assesses presentations slide by slide, matching human system rankings and detecting issues such as falsified numbers, degraded figures, dropped slides, and shuffled slide order.
arXiv:2606. 02080v1 Announce Type: cross Abstract: Biological image analysis increasingly demands integration across heterogeneous tools, programming environments, and domain knowledge that few researchers can command simultaneously.
arXiv:2607. 21946v1 Announce Type: new Abstract: This technical report presents our approach to Challenge Track~3: SeePhys Pro at the 3rd AI for Math Workshop, where the task is to answer college-level physics questions whose statement and figure may be given partly or entirely as an image.