arXiv AI

FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures

arXiv AI
Aug 24

MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature

MatMMExtract is an open‑source pipeline that disassembles compound scientific figures into individual sub‑panels and generates structured, grounded image‑text pairs using a large language model guided by a materials science taxonomy. Applied to 14,810 open‑access articles, it produced 391,606 panel‑level pairs with sub‑captions, a two‑level visualisation category (19 classes, 100+ subtypes), and scientific summaries. The project also introduces MaterialScope, a 2,811‑figure detection dataset, and demonstrates that Gemini 3.1 Flash Lite yields high‑quality annotations with low hallucination, while a dual‑encoder baseline outperforms zero‑shot CLIP on the resulting MatSciFig dataset.

By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
Hugging Face Trending Papers
Jul 29

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers.

arXiv AI
Jul 31

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

arXiv:2607. 27066v1 Announce Type: cross Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy.

By Chuanzhi Xu, Zihan Deng, Huiqi Liang, Chengkun Yue, Zhanlin Cui, Pengfei Ye, Weidong Cai
arXiv AI
4d ago

FigAct: Turning Scientific Figures into Active Canvases for Explanation

FigAct transforms static scientific figures into question‑conditioned visual presentations by acting directly on existing graphical elements. The framework generates short narrations, grounds each narration in visual evidence, and applies visual actions to guide viewer attention, mimicking a human presenter. A hierarchical search strategy reduces token usage by about 40×, and FigAct‑8B is trained with rewards for grounding accuracy, search efficiency, and rendering quality, evaluated on a human‑verified benchmark of real‑world scientific figures.

By Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey
arXiv AI
Aug 3

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

arXiv:2607. 28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and frozen pretrained image-text scorers.

By Duy Tran Thanh, Thien-Phuc Doan, Long Nguyen-Vu, Ngo Tan Vu Khanh