arXiv:2604. 09552v2 Announce Type: replace-cross Abstract: Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retrieval augmented generation (RAG) systems.
By Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
AiSearch is a flexible multimodal retrieval framework that uses Vision Language Models (VLMs) to enable natural language search over images and videos. It supports interactive search refinement through user feedback, allowing results to be tailored to the user's intent in real time. The system also provides visual benchmarking across multiple VLMs, enabling users to choose the most suitable model for their specific task.
By Ali Koksal, Mei Chee Leong, Vicky Sintunata, Ching Ling Chin, Wee Teck Fong
The paper presents a semester-long deployment of the VideoPoints platform, featuring a retrieval‑augmented chatbot that answers questions using only the active course’s lecture materials and provides clickable, timestamped citations. Across 833 interactions, 70.5% of messages included citations, no cross‑course references were made, and the bot declined to answer when no lecture evidence matched. The study also shows that the system improves correct‑lecture retrieval by 6.3 percentage points over dense‑only retrieval on the EduVidQA benchmark.
By S M Masrur Ahmed, Jaspal Subhlok
The paper introduces TrioRAG, a graph-free multimodal retrieval-augmented generation framework that combines evidence from the question, an anchor image, and a VLM-enhanced query via late fusion. It also presents AutoQA, a benchmark featuring noisy web-sourced images that require reasoning across manuals. TrioRAG outperforms graph-based systems on three benchmarks while cutting costs and speeding up inference by 1.6–2.3×.
By Tithi Rakshit, Hongkuan Zhou, Lavdim Halilaj, Yuqicheng Zhu
arXiv:2508.13186v2 Announce Type: replace-cross
Abstract: AI agents with advanced reasoning and tool-use capabilities have demonstrated impressive performance in web browsing for deep search. However...
By Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, Chuanhao Li, Jiayi Tian, Chenchen Zhang, Tianhao Peng, Yancheng He, Jihao Gu, Hui Huang, Donghao Zhou, Yuanxing Zhang, Jian Yang, Ge Zhang, Wenhao Huang, Zhaoxiang Zhang, Qiangpeng Yang, Shilei Wen
arXiv:2607. 08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing.
By Feng Wang, Canmiao Fu, Zhipeng Huang, Chen Li, Jing Lyu, Ge Li