Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 15753v1 Announce Type: new Abstract: Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual grounding throughout multi-step reasoning.
arXiv:2507.00754v3 Announce Type: replace Abstract: The integration of Large Language Model (LLMs) blocks with Vision Transformers (ViTs) holds immense promise for vision-only tasks by leveraging the...
arXiv:2602. 14065v2 Announce Type: replace Abstract: Knowledge-intensive Visual Question Answering (KI-VQA) frequently suffers from severe knowledge conflicts caused by the inherent limitations of open-domain retrieval.
WeakMCN introduces a multi-task collaborative network that jointly learns weakly supervised referring expression comprehension (WREC) and segmentation (WRES) using a dual-branch architecture. The WREC branch employs anchor-based contrastive learning and serves as a teacher for the WRES branch, while two novel modules—Dynamic Visual Feature Enhancement (DVFE) and Collaborative Consistency Module (CCM)—facilitate cross-task collaboration. Experiments on RefCOCO, RefCOCO+, and RefCOCOg show significant performance gains over single-task baselines, with up to 3.91% and 13.11% improvements on WREC and WRES respectively, and strong generalization in semi-supervised settings.
arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.
arXiv:2606. 30498v1 Announce Type: cross Abstract: Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color.