The paper investigates how large language models (LLMs) interpret ambiguous or incomplete text prompts for visualization authoring and introduces visual prompts as a complementary modality to improve precision. An empirical study informs the design of VisPilot, a system that allows users to create visualizations using text, sketches, and direct manipulation. A controlled user study and expert evaluation show that multimodal prompts help users convey spatial constraints, local references, and design preferences while maintaining task efficiency comparable to text-only prompting.
By Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, Wei Chen
Editable Visual Design introduces a new design paradigm that combines a Coding Agent with a Vision‑Language Model (VLM) and an image generation model. The VLM acts as the creative brain, understanding requirements, planning tasks, and judging aesthetics, while the image generator produces isolated visual assets on demand. The agent follows an "imagine first, then act" workflow, generating assets, writing native HTML/CSS, and refining the design through visual feedback, ultimately producing editable, layer‑wise artifacts with real text that can be adjusted via a graphical interface.
By Junyan Ye, Wei Liu, Dongzhi Jiang, Zichen Wen, HaoDong Li, Zhutao Lv, Jiaxin Lin, Jinhua Yu, Jun He, Zilong Huang, Rui Chen, Weijia Li
The paper introduces CanvasConvo, a system that presents large language model (LLM) conversations in two synchronized views: a traditional linear chat for ongoing dialogue and a spatial canvas that visualizes the conversation’s branching structure. In a five‑day field study with 24 participants, users tended to switch between the views rather than replace chat entirely; chat remained the primary interaction mode while the canvas was used for overview, revisiting, and exploring alternative paths. The results highlight challenges such as entrenched chat habits, smooth transitions between representations, and understanding branch context, offering insights for designing LLM interfaces that blend linear and non‑linear conversation representations.
By Rifat Mehreen Amin, Alperen Adatepe, Daniela Fernandes, Daniel Buschek, Andreas Butz
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.
By Jackson Vonderhorst, Kuangshi Ai, Haichao Miao, Shusen Liu, Chaoli Wang
PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.
By Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng
Teaching machines to emulate natural handwriting styles remains an open challenge, as it requires synthesizing stroke sequences that dynamically vary in shape, texture, pressure and script - not only across individuals, but also within a single person's handwriting. Attempts at this challenge have largely explored deep learning methods in both online and offline settings.