arXiv AI

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

arXiv:2607. 18514v1 Announce Type: cross Abstract: Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text.

arXiv AI
Aug 17

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

arXiv:2608. 14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret.

By Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden
arXiv AI
Aug 24

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

VIALS is a new visual question‑answering benchmark comprising 161 interpretation tasks based on real scientific artifacts such as gel blots, microscopy images, plasmid maps, flow cytometry plots, and molecular structures. The benchmark focuses on the types of images encountered in everyday biotech workflows rather than polished figures from publications. Current state‑of‑the‑art vision‑language models struggle to interpret these domain‑specific images, whereas scientists with relevant expertise find the tasks straightforward.

By Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor, Francisco Guzm\'an, Nicholas Magazine, Jonas Mueller
arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas
arXiv AI
Jun 18

Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.

By Ayae Ide, Tory Park, Jaron Mink, Tanusree Sharma
arXiv AI
Sep 12

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

The paper investigates how large language models (LLMs) interpret ambiguous or incomplete text prompts for visualization authoring and introduces visual prompts as a complementary modality to improve precision. An empirical study informs the design of VisPilot, a system that allows users to create visualizations using text, sketches, and direct manipulation. A controlled user study and expert evaluation show that multimodal prompts help users convey spatial constraints, local references, and design preferences while maintaining task efficiency comparable to text-only prompting.

By Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, Wei Chen
arXiv AI
Sep 15

NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

NoteVQA is a new benchmark that collects 252 real‑life visual questions from the Chinese image‑sharing platform Xiaohongshu, covering 12 topics and 7 user intents. Each question is paired with a concise expert reference and a human‑audited interleaved answer that blends text and visual evidence. The study evaluates VLMs on short‑answer correctness and interleaved answer quality using a new AgenticInterleave framework and a 12‑dimensional IVR‑12 rubric, finding that even state‑of‑the‑art models achieve only about 53% accuracy and lag behind human references in content quality.

By Haonan Jiang, Guojian Zhan, Jiancong Xie, Shijun Wan, Dongiia Zhao, Cheng Chen, Yahui Liu, Yao Hu, Chuan Mu
arXiv AI
Sep 2

MIDR: Enrichment-Augmented Indexing for Multimodal Document Retrieval

MIDR (Multimodal Indexing for Document Retrieval) is a training‑free framework that enriches document indexes by converting rendered pages into verified textual fields with a multimodal LLM, then indexing those fields with BM25F and optionally fusing with dense retrieval. By shifting multimodal reasoning to index time, MIDR enables text‑centric serving while retaining multimodal evidence, achieving a 23.0% relative gain over BM25 on ViDoRe V3 and outperforming ColQwen2.5 on several domains with significantly smaller index memory and lower query latency.

By Debanjan Mahata, Atharva Tendle, Daniel Preotiuc-Pietro, Yong Zhuang, Ozan Irsoy
arXiv AI
4d ago

FigAct: Turning Scientific Figures into Active Canvases for Explanation

FigAct transforms static scientific figures into question‑conditioned visual presentations by acting directly on existing graphical elements. The framework generates short narrations, grounds each narration in visual evidence, and applies visual actions to guide viewer attention, mimicking a human presenter. A hierarchical search strategy reduces token usage by about 40×, and FigAct‑8B is trained with rewards for grounding accuracy, search efficiency, and rendering quality, evaluated on a human‑verified benchmark of real‑world scientific figures.

By Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey
arXiv AI
Sep 2

SCAFFOLD: A Large-Scale Structured Dataset of Computer Science Research Figures with Diagram QA and Chain-of-Thought Reasoning Traces

SCAFFOLD is a large-scale structured dataset of computer science research figures, each paired with captions, context, questions, answers, and chain-of-thought reasoning traces. It contains 157,387 figure–question pairs from 3,058 arXiv papers, with subsets of 36,797 and 12,000 pairs for medium and small-scale use. The dataset was created using layout detection, PDF parsing, and AI-assisted question generation, and was used to benchmark a vision‑language model (Qwen2.5‑VL‑3B‑Instruct).

By Ranjit Raut, Aarav Subedi, Sagun Rai, Sudan Jha
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso