Visual Abstention in Unified Multimodal Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2610.03002v1 Announce Type: new Abstract: Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvem...
arXiv:2608. 07565v1 Announce Type: cross Abstract: Conversational assistants increasingly recommend follow-up edits to help users continue a task.
The paper introduces an interventional protocol to assess how vision‑language models (VLMs) explain the impact of missing modalities on their predictions. By comparing the models’ self‑explanations with actual changes observed after restoring missing inputs, the study finds that VLMs routinely overstate the sufficiency of available evidence and underestimate the effect of adding back missing modalities. Across eight open‑weight VLMs and four tasks, the discrepancy between predicted and realized changes is substantial, revealing systematic mischaracterization of modality dependence.
arXiv:2606. 01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction following, minimally editing the source image, and ensuring high visual quality.
Aphanta is an automated framework that diagnoses how well image editors can produce task‑aligned visual intermediates for multimodal large language models (MLLMs). It evaluates three reasoning conditions—direct, editor‑generated, and idealized intermediate—to distinguish visual potential from practical editor performance across 20 tasks and various editor–MLLM pairs. The study finds that image editing benefits certain tasks like visual cue injection and grounding, but is less reliable for symbol‑sensitive or structural tasks, and demonstrates measurable performance gains with a Qwen pipeline.
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles.