FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
MatMMExtract is an open‑source pipeline that disassembles compound scientific figures into individual sub‑panels and generates structured, grounded image‑text pairs using a large language model guided by a materials science taxonomy. Applied to 14,810 open‑access articles, it produced 391,606 panel‑level pairs with sub‑captions, a two‑level visualisation category (19 classes, 100+ subtypes), and scientific summaries. The project also introduces MaterialScope, a 2,811‑figure detection dataset, and demonstrates that Gemini 3.1 Flash Lite yields high‑quality annotations with low hallucination, while a dual‑encoder baseline outperforms zero‑shot CLIP on the resulting MatSciFig dataset.
arXiv:2606. 29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale.
arXiv:2607. 27084v1 Announce Type: cross Abstract: Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers.
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers.
arXiv:2610.00952v1 Announce Type: cross Abstract: Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-t...
arXiv:2609.00709v1 Announce Type: cross Abstract: Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions sho...