arXiv Computation and Language
Sep 18

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

The paper introduces a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, targeting both the understanding (SIQA-U) and scoring (SIQA-S) tracks of the SIQA challenge. It builds a multimodal index that merges textual semantics with fine‑grained visual features and employs a multi‑route retrieval and fusion mechanism to supply large language models with relevant reference cases, improving their evaluation of complex scientific images. The approach aligns well with human expert judgment and secured first place in the SIQA-U track at the ICME 2026 Grand Challenges.

By Yinuo Zhang, Bingshuo Liu, Zhiying Tu, Dianhui Chu, Qingbin Liu, Xi Chen, Jiang Bian, Xiaoyan Yu, Dianbo Sui
Hugging Face Trending Papers
Jul 29

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers.

arXiv Computation and Language
Sep 17

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

MiRAGE is a new evaluation framework for retrieval‑augmented generation (RAG) that handles multimodal sources such as audiovisual media. It uses a claim‑centric approach with two metrics: InfoF1, which measures factuality and information coverage, and CiteF1, which measures citation support and completeness. Human evaluation shows MiRAGE aligns well with extrinsic quality judgments, and an automatic implementation outperforms three text‑centric RAG metrics (ALCE, ARGUE, RAGAS) on text while uniquely generalizing to multimodal inputs.

By Alexander Martin, William Walden, Reno Kriz, Dengjia Zhang, Kate Sanders, Eugene Yang, Chihsheng Jin, Benjamin Van Durme
arXiv AI
Jun 10

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

arXiv:2606. 10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models.

By Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan, Muhammad Haris Khan
arXiv AI
Aug 17

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

arXiv:2608. 14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret.

By Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden