arXiv AI By Shenghui Chen, Po-han Li, Ximeng Sun, Shijia Yang, Emad Barsoum, Zicheng Liu, Sandeep Chinchali, Ufuk Topcu

VEGAS: Human-Aligned Video Caption Evaluation via Gaze

Read the original on arXiv AI →

arXiv:2607. 08489v1 Announce Type: cross Abstract: Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

CapMem is a new benchmark for evaluating caption-based episodic memory in egocentric video. It contains 75 videos (33.7 hours total) and 1,000 multiple-choice questions across 16 scenarios, designed to test the Episodic Memory Video Caption QA task. Experiments show that using captions as memory outperforms direct VideoQA on long videos, and a caption-guided retrieve-and-verify approach further boosts accuracy.

By Dingli Liang, Yiqiao Xie, Yukai Huang, Zhaokai Wang, Weitong Cai, Guangwen Feng, Jifei Song, Zhensong Zhang, Hang Zhang
Hugging Face Trending Papers
Jun 29

Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation

Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgments. Recent approaches using large language models (LLMs), commonly referred to as LLM-as-a-Judge, have improved alignment with human judgments but still suffer from a mismatch between large-vocabulary language modeling and evaluation over a small label set.

Hugging Face Trending Papers
Jul 23

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.