arXiv Computer Vision By Chujie Qin, Zilong Zhang, Zewei Chang, Chunle Guo, Ruixing Wang, Tao Hu, Ming-Ming Cheng, Chongyi Li

Dotting the Eye: An Intent-Driven Image Retouching Agent for Visual Focus Enhancement

Read the original on arXiv Computer Vision →

The paper introduces EyeControl, an intent-driven image retouching agent that enhances visual focus by guiding attention to a target region with minimal user input. It combines a multi‑modal large language model to interpret user intent and a diffusion‑based retouching executor that aligns its attention map with a pseudo‑intent map, while an operation‑consistency constraint ensures natural global and local adjustments. The authors also present ControlArt‑Bench, a dataset for evaluating visual focus enhancement, and demonstrate that EyeControl achieves perceptually appealing results with stronger intent alignment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 3

See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

arXiv:2606. 02735v1 Announce Type: cross Abstract: Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control.

By Yueh-Hua Wu, Tatsuya Matsushima, Kei Ota
arXiv AI
Jun 9

IEA: Amateur-Friendly Conversational Image Editing Agent via Three Stages of Multitask Alignment

arXiv:2606. 08016v1 Announce Type: cross Abstract: Current image editing software often hinges on fixed filters or expert tuning, leaving a gap between amateur users' intent and outcomes.

By Zichen Zhu, Yuheng Sun, Mingxuan Zhu, Wenjie Ma, Situo Zhang, Zhexiang Wang, Ziyue Yang, Danyang Zhang, Kunyao Lan, Zihan Zhao, Dingye Liu, Siqi Xiang, Lu Chen, Kai Yu