arXiv Computer Vision

Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance

The paper introduces Salience-LLaVA, a vision‑language model that prioritizes scene elements based on their importance for low‑vision users. It presents three new salience‑aware datasets—Salience COCO, Salience Flickr, and Salience VizWiz—annotated with object‑level salience verified by low‑vision participants. The authors also propose the SCMI metric to evaluate caption ordering accuracy and demonstrate the system’s practicality by deploying it on assistive glasses.

Hugging Face Trending Papers
Jul 23

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.

arXiv Computer Vision
2d ago

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

arXiv:2609.00591v1 Announce Type: new Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...

By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon