arXiv Computer Vision

Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement

The paper introduces EyeControl, an intent-driven image retouching agent that enhances visual focus by guiding attention to a target region with minimal user input. It combines a multi‑modal large language model to interpret user intent and a diffusion‑based retouching executor that aligns its attention map with a pseudo‑intent map, while an operation‑consistency constraint ensures natural global and local adjustments. The authors also present ControlArt‑Bench, a dataset for evaluating visual focus enhancement, and demonstrate that EyeControl achieves perceptually appealing results with stronger intent alignment.

arXiv AI
Jun 3

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.

By Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
arXiv AI
Jun 9

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

arXiv:2606. 08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes.

By Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu
arXiv AI
Sep 7

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

The paper investigates how visuomotor imitation policies fail when visually similar distractors are introduced, framing the issue as conditional visual grounding where the target changes with manipulation phase and task state. Using Action Chunking with Transformers (ACT), the authors systematically vary color and shape similarity of distractors, pinpointing failures to picking and placement stages. They then test distractor augmentation, phase‑dependent attention regularization, and appearance‑based visual prompting, which together significantly improve robustness in simulation and on a physical UR3e robot, and demonstrate similar improvements on a pretrained vision‑language‑action policy for instrument handling.

By Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger