arXiv Computation and Language By Yinuo Zhang, Bingshuo Liu, Zhiying Tu, Dianhui Chu, Qingbin Liu, Xi Chen, Jiang Bian, Xiaoyan Yu, Dianbo Sui

Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

Read the original on arXiv Computation and Language →

The paper introduces a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, targeting both the understanding (SIQA-U) and scoring (SIQA-S) tracks of the SIQA challenge. It builds a multimodal index that merges textual semantics with fine‑grained visual features and employs a multi‑route retrieval and fusion mechanism to supply large language models with relevant reference cases, improving their evaluation of complex scientific images. The approach aligns well with human expert judgment and secured first place in the SIQA-U track at the ICME 2026 Grand Challenges.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Jul 29

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers.

arXiv Computation and Language
Sep 17

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

MiRAGE is a new evaluation framework for retrieval‑augmented generation (RAG) that handles multimodal sources such as audiovisual media. It uses a claim‑centric approach with two metrics: InfoF1, which measures factuality and information coverage, and CiteF1, which measures citation support and completeness. Human evaluation shows MiRAGE aligns well with extrinsic quality judgments, and an automatic implementation outperforms three text‑centric RAG metrics (ALCE, ARGUE, RAGAS) on text while uniquely generalizing to multimodal inputs.

By Alexander Martin, William Walden, Reno Kriz, Dengjia Zhang, Kate Sanders, Eugene Yang, Chihsheng Jin, Benjamin Van Durme
arXiv AI
Aug 17

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

arXiv:2608. 14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret.

By Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden