arXiv Computer Vision By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati

OpenVAM: Open-World Visual Attention Modeling with VLMs

Read the original on arXiv Computer Vision →

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
4d ago

Amplify What You Gaze At: Target Saliency Boosting in Text-to-Image Generation

The paper introduces Target Saliency Boosting, a new task that enhances the visual prominence of a specific object in text-to-image generation without visual priors. It proposes GazeME, a lightweight framework that inserts learnable marker tokens around object descriptions to indicate which objects to emphasize or suppress. By building a saliency-semantics dataset and using Saliency Prior Marker Activation, GazeME learns to adjust markers during training and automatically applies them at inference, effectively boosting target saliency while maintaining semantic alignment and image quality.

By Shengqi Dang, Zhengxi Yu, Feilin Han, Xingyu Lan, Nan Cao
arXiv Computer Vision
Aug 31

Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance

The paper introduces Salience-LLaVA, a vision‑language model that prioritizes scene elements based on their importance for low‑vision users. It presents three new salience‑aware datasets—Salience COCO, Salience Flickr, and Salience VizWiz—annotated with object‑level salience verified by low‑vision participants. The authors also propose the SCMI metric to evaluate caption ordering accuracy and demonstrate the system’s practicality by deploying it on assistive glasses.

By Jiazhao Liang, Hao Huang, Shuaihang Yuan, Congcong Wen, Geeta Chandra Raju Bethala, Giles Hamilton-Fletcher, Yu Hao, John-Ross Rizzo, Mengyu Wang, Anthony Tzes, Yi Fang