arXiv:2608.21431v1 Announce Type: new
Abstract: Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information...
By Qiyou Liu, Yong Zhang, Jianjie Luo, Zhenguo Yang, Yi Yu
arXiv:2604.07146v3 Announce Type: replace
Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for...
By Zhuohong Chen, Zhenxian Wu, Yunyao Yu, Hangrui Xu, Zirui Liao, Zhifang Liu, Xiangwen Deng, Pen Jiao, Haoqian Wang
arXiv:2608.21796v1 Announce Type: cross
Abstract: Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond the visua...
By Long Shu, Shuochen Liu, Wei Chen, Junda Lin, Zhi Zheng, Huijun Hou, Tong Xu
arXiv:2507. 20804v3 Announce Type: replace Abstract: Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge.
By Xueyao Wan, Hang Yu
arXiv:2607. 22643v1 Announce Type: new Abstract: Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still retrieve directly from raw multimodal input over a flat evidence space.
By Tianyu Yang, Shir Simon, Zhenzhen Li, Minhao Cheng, Xiangliang Zhang
arXiv:2608.29644v1 Announce Type: cross
Abstract: Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art his...
By Marc S. Walton, Astrid Harth
arXiv:2606. 31222v1 Announce Type: new Abstract: Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction.
By Gunho Jung, Jeong-Woo Park, Seon Bin Kim, Seong-Whan Lee
arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.
By Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
arXiv:2505.03654v3 Announce Type: replace-cross
Abstract: Multimodal Large Language Models have shown strong performance across multimodal tasks, and recent personalized MLLMs can recognize user-spec...
By Yifan Xiang, Zhenxi Zhang, Bin Li, Yixuan Weng, Bo Gao, Shoujun Zhou, Yangfan He, Yilin Yuan, Keqin Li
arXiv:2608.21450v1 Announce Type: new
Abstract: Knowledge-Based Visual Question Answering (KB-VQA) relies on retrieving external information to answer queries involving long-tail entities. However, e...
By Hangrui Xu, Zhengxian Wu, Yunyao Yu, Zhuohong Chen, Rui Cong, Xiangwen Deng, Zhifang Liu, Peng Jiao, Haoqian Wang
Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-augmented image captioning framework that generates captions with deeper insights, such as object attributes, event context, and underlying significance, by leveraging external knowledge.
arXiv:2503. 10677v3 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language understanding and generation by combining large-scale retrieval systems with generative models.
By Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, Daoyu Wang, Enhong Chen