VEGAS: Human-Aligned Video Caption Evaluation via Gaze
arXiv:2607. 08489v1 Announce Type: cross Abstract: Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention.
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose VEGAS (Video caption Evaluation via GAze Score), a training-free metric that leverages test-time gaze to sample personalized, attention-aligned text.
arXiv:2607. 08489v1 Announce Type: cross Abstract: Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention.
arXiv:2609.09973v1 Announce Type: new Abstract: Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generate...
CapMem is a new benchmark for evaluating caption-based episodic memory in egocentric video. It contains 75 videos (33.7 hours total) and 1,000 multiple-choice questions across 16 scenarios, designed to test the Episodic Memory Video Caption QA task. Experiments show that using captions as memory outperforms direct VideoQA on long videos, and a caption-guided retrieve-and-verify approach further boosts accuracy.
Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgments. Recent approaches using large language models (LLMs), commonly referred to as LLM-as-a-Judge, have improved alignment with human judgments but still suffer from a mismatch between large-vocabulary language modeling and evaluation over a small label set.
Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.
arXiv:2606.29997v2 Announce Type: replace Abstract: Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limi...
arXiv:2603.04349v2 Announce Type: replace Abstract: Understanding long videos is crucial for embodied intelligent agents, as their performance depends on effectively accumulating and using long-horiz...
VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.
arXiv:2511. 19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external supervision.
arXiv:2308. 06035v4 Announce Type: replace Abstract: Humans routinely draw on visual context to predict upcoming words.
arXiv:2607. 17994v1 Announce Type: cross Abstract: Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation.
The paper introduces Salience-LLaVA, a vision‑language model that prioritizes scene elements based on their importance for low‑vision users. It presents three new salience‑aware datasets—Salience COCO, Salience Flickr, and Salience VizWiz—annotated with object‑level salience verified by low‑vision participants. The authors also propose the SCMI metric to evaluate caption ordering accuracy and demonstrate the system’s practicality by deploying it on assistive glasses.