Agentic Authoring of Interactive Multiview Visualizations in Genomics
arXiv:2606. 00370v1 Announce Type: cross Abstract: Diverse genomics data, scientific questions, and analysis tasks typically demand highly specialized visualizations.
The paper investigates how large language models (LLMs) interpret ambiguous or incomplete text prompts for visualization authoring and introduces visual prompts as a complementary modality to improve precision. An empirical study informs the design of VisPilot, a system that allows users to create visualizations using text, sketches, and direct manipulation. A controlled user study and expert evaluation show that multimodal prompts help users convey spatial constraints, local references, and design preferences while maintaining task efficiency comparable to text-only prompting.
arXiv:2606. 00370v1 Announce Type: cross Abstract: Diverse genomics data, scientific questions, and analysis tasks typically demand highly specialized visualizations.
arXiv:2609.26208v1 Announce Type: new Abstract: Data visualization is central to analytical reasoning, but real-world analysis increasingly requires language-driven interactive interfaces rather than...
arXiv:2606. 08492v1 Announce Type: cross Abstract: Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity of user prompts.
arXiv:2606. 26614v1 Announce Type: cross Abstract: Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis).
arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.
arXiv:2607. 25911v1 Announce Type: cross Abstract: Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints.
arXiv:2505. 04260v3 Announce Type: replace-cross Abstract: Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold start and difficult to articulate in natural language.
The paper surveys multimodal speculative decoding, examining whether diffusion-based block‑parallel generative drafting—successful in text‑only LLMs—can be applied to Vision‑Language, Video‑Language, Audio, and Vision‑Language‑Action models. It introduces a taxonomy separating drafter‑side parallelism from other design choices, and presents an empirical comparison across benchmarks such as OCR, VQA, visual reasoning, and image captioning. The study highlights current limitations, outlines open challenges, and suggests future research directions for multimodal speculative decoding.
arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...
arXiv:2510. 04514v3 Announce Type: replace Abstract: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts-those requiring precise visual interpretation rather than relying on textual shortcuts.
arXiv:2607. 18514v1 Announce Type: cross Abstract: Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text.
arXiv:2604. 27996v3 Announce Type: replace Abstract: This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visualization workflows from natural-language instructions.