arXiv:2609.00591v1 Announce Type: new
Abstract: An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level c...
By Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi, Franck Dernoncourt, Jack Wang, Koustava Goswami, Nedim Lipka, Puneet Mathur, Samyadeep Basu, Seunghyun Yoon
arXiv:2608. 19825v1 Announce Type: cross Abstract: Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems.
By Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim
VidOmni-Bench is a new benchmark for fine‑grained video understanding that asks models to verify whether each event in dense video captions is supported by the video. It contains 500 videos covering five complexity types and durations from 4 seconds to 90 minutes, and uses human‑verified sentence‑level labels to create hard negatives. Experiments show that Video‑LLMs often hallucinate events, struggle to detect incorrect descriptions, and exhibit varying weaknesses depending on video complexity and duration.
By Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, Donghyun Kim
arXiv:2607. 09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models.
By Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu
MedGEN-Bench is a new benchmark for open‑ended multimodal medical generation that addresses limitations in current medical visual benchmarks, such as query‑image misalignment, closed‑ended answer spaces, and text‑centric outputs. The dataset contains 6,422 image‑text pairs across six imaging modalities, 15 clinical tasks, and 27 subtasks, including VQA, image editing, and contextual multimodal generation pairs. Evaluation combines reference‑based fidelity metrics with a structured, checklist‑guided assessment by a medical VLM judge, and preliminary results show that image‑output tasks remain unsaturated while contextual augmentation improves image‑instruction similarity.
By Junjie Yang, Yuhao Yan, Gang Wu, Rui Qian, Zhisheng Chen, Haijiang Li, Yuhe Wu, Qichao Zhao, Dawen Tian, Xiang Wan, Fenglei Fan, Wenjian Qin, Yongquan Zhang, Feiwei Qin, Changmiao Wang
Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visua...
arXiv:2606. 11925v1 Announce Type: cross Abstract: Sign language translation (SLT) converts sign language video into spoken language text and holds significant promise for improving accessibility and enabling communication between signing and non-signing communities.
By Zsolt Robotka, \'Ad\'am R\'ak, Jalal Al-Afandi, Andr\'as Horv\'ath, Gy\"orgy Cserey
VTR-Bench is a new benchmark designed to evaluate how well video generation models render text within scenes. It includes 300 prompts across five real-world scenarios such as advertisements and scientific videos, and uses an automated pipeline with human alignment to assess text fidelity and scene/motion requirements. Experiments on 11 state‑of‑the‑art models show that even the best performer has a word error rate of 0.250, underscoring widespread challenges in visual text rendering.
By Yu Huang, Jungang Li, Zhiyuan Wang, Yonghua Hei, Song Dai, Jiayu Yang, Deyuan Liu, Xiang Zheng, Xiaoshuang Shi, Hao Cheng, Kaidi Xu
The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.
By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach
Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.
arXiv:2603.22179v2 Announce Type: replace
Abstract: Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Curr...
By Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari, Tahoura Nedaee, Francois Haddad, Ivan Lopez, Fang Cao, Michael Salerno, Li Fe-Fei, Ehsan Adeli, Rima Arnaout, Euan A Ashley
The paper introduces Redemption Score (RS), a multi‑modal evaluation framework for image captioning that combines three complementary signals: Mutual Information Divergence for global image‑text alignment, DINO‑based perceptual similarity of cycle‑generated images for visual grounding, and LLM text embeddings for contextual similarity to human references. RS fuses these signals to provide a more holistic assessment, achieving a Kendall‑τ of 58.42 on Flickr8k and outperforming most prior methods. The framework demonstrates consistent performance across Conceptual Captions and MS COCO, offering a robust evaluation that captures both visual accuracy and text quality.
By Ashim Dahal, Ankit Ghimire, Saydul Akbar Murad, Nick Rahimi