arXiv AI By Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw, Jennifer Healey

FigAct: Turning Scientific Figures into Active Canvases for Explanation

Read the original on arXiv AI →

FigAct transforms static scientific figures into question‑conditioned visual presentations by acting directly on existing graphical elements. The framework generates short narrations, grounds each narration in visual evidence, and applies visual actions to guide viewer attention, mimicking a human presenter. A hierarchical search strategy reduces token usage by about 40×, and FigAct‑8B is trained with rewards for grounding accuracy, search efficiency, and rendering quality, evaluated on a human‑verified benchmark of real‑world scientific figures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

The paper introduces SciGram, a large-scale dataset of 194K scientific diagrams paired with 1.4M visual instructions generated through a terminology‑grounded pipeline that extracts domain concepts, synthesizes facts, and retrieves relevant diagrams. Models fine‑tuned on SciGram show significant gains on diagram‑centric benchmarks such as TQA, ScienceQA, and AI2D, and when combined with existing models like LLaVA OneVision, set new state‑of‑the‑art performance. The authors release both the dataset and trained models to support further research in scientific diagram understanding.

By Raul Ortega, Jos\'e Manuel G\'omez-P\'erez
arXiv AI
Sep 3

TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement Learning

TikZilla is a new approach to generating TikZ code from textual descriptions, built on a larger, higher‑quality dataset called DaTikZ‑V4 that includes LLM‑generated figure descriptions. The method uses a two‑stage pipeline: supervised fine‑tuning of small Qwen models (3B and 8B) followed by reinforcement learning with an image encoder that provides semantically faithful reward signals. Human evaluations show that TikZilla outperforms its base models by 1.5–2 points on a 5‑point scale, beats GPT‑4o by 0.5 points, and matches GPT‑5 in image‑based tests while remaining much smaller.

By Christian Greisinger, Steffen Eger
arXiv AI
Sep 2

Figures as Programs: Recursive Generation of Editable Scientific Figures

The paper introduces “FigTree”, a multi-agent system that automatically converts a scientific paper into a structured vector figure by recursively constructing SVG programs. It decomposes figures into hierarchical regions, generates each region as a short SVG program, and assembles them, using a render‑critic refinement loop to trace and repair visual defects. Evaluations show that “FigTree” produces high‑quality figures and allows more effective editing than raster‑based methods.

By Yepeng Liu, Dasen Dai, Chengzhi Liu, Yiren Song, Hai Ci, Yu Zhang, Qi Zhang, Mike Zheng Shou, Xin Eric Wang, Yuheng Bu
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso