arXiv AI

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

The paper investigates how large language models (LLMs) interpret ambiguous or incomplete text prompts for visualization authoring and introduces visual prompts as a complementary modality to improve precision. An empirical study informs the design of VisPilot, a system that allows users to create visualizations using text, sketches, and direct manipulation. A controlled user study and expert evaluation show that multimodal prompts help users convey spatial constraints, local references, and design preferences while maintaining task efficiency comparable to text-only prompting.

arXiv AI
Jun 11

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.

By Jana Zeller, Thadd\"aus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell, Wieland Brendel
arXiv AI
Aug 24

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

The paper surveys multimodal speculative decoding, examining whether diffusion-based block‑parallel generative drafting—successful in text‑only LLMs—can be applied to Vision‑Language, Video‑Language, Audio, and Vision‑Language‑Action models. It introduces a taxonomy separating drafter‑side parallelism from other design choices, and presents an empirical comparison across benchmarks such as OCR, VQA, visual reasoning, and image captioning. The study highlights current limitations, outlines open challenges, and suggests future research directions for multimodal speculative decoding.

By Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan, Qiang Hui, Shuting Liu, Fuyuan Shi, Ting Lu, Shaoan Zhao, Xueqiang Guo, Xinpei Su, Jianbing Zhang, Xinyu Dai, Kai Wang, Shiguo Lian
arXiv AI
Sep 15

ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation

arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...

By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
arXiv AI
Jun 10

ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering

arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.

By Rachneet Kaur, Nishan Srishankar, Zhen Zeng, Sumitra Ganesh, Manuela Veloso