Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,971 stories · RSS feed

arXiv Computer Vision
Sep 30

From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection

arXiv:2609.36145v1 Announce Type: new Abstract: Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective...

By Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo, Xiaoye Qiu, Haodong Li, Taiping Yao, Bo Wang, Youchang Xiao, Bin Li, Shouhong Ding
arXiv AI
Sep 30

ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.

By Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao
arXiv AI
Sep 30

Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination

The paper introduces Embodied Semantic Communication (ESC), a new paradigm that redefines information transmission for autonomous agents by embedding multimodal perceptual states, hardware capabilities, and collaborative intents into unified, action‑oriented semantic representations. ESC enables heterogeneous agents to parse, align, and ground shared information directly into local motor control, addressing the limitations of traditional communication approaches that focus solely on bit delivery or single‑task optimization. The tutorial outlines ESC’s conceptual boundaries, system characteristics, and technical pathways, mapping relevant mathematical tools such as semantic information theory, world models, and multi‑agent decision theory, and concludes with a roadmap of open challenges like semantic reliability, dynamic interaction, and bandwidth‑adaptive transmission.

By Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du, Wenchi Cheng, Zhu Han
arXiv AI
Sep 30

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.

By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu