arXiv AI

Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

arXiv Computer Vision
6d ago

WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation

WeakMCN introduces a multi-task collaborative network that jointly learns weakly supervised referring expression comprehension (WREC) and segmentation (WRES) using a dual-branch architecture. The WREC branch employs anchor-based contrastive learning and serves as a teacher for the WRES branch, while two novel modules—Dynamic Visual Feature Enhancement (DVFE) and Collaborative Consistency Module (CCM)—facilitate cross-task collaboration. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show significant performance gains over single-task baselines, with up to 3.91% and 13.11% improvements on WREC and WRES respectively, and strong generalization in semi-supervised settings.

By Silin Cheng, Yang Liu, Xinwei He, Sebastien Ourselin, Lei Tan, Gen Luo
arXiv Computer Vision
Aug 21

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.

By Yanxiang Huang, Guohua Gao, Zhaoyang Wei
arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang