arXiv:2607. 29440v1 Announce Type: new Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions.
By Zhoujin Tian, Yao Tian, Hao Zhang, Cheng Chen, Yakun Li, Lei Zhang, Xiaofang Zhou
arXiv:2608.29897v1 Announce Type: new
Abstract: Long-horizon agents need a context manager to compress growing interaction histories into a bounded working context, via passive strategies or active s...
By Jiaqi Su, Cong Pang, Jiawei Hong, Tiankuo Yao, Zixuan Chen, Xin Lou, Lewei Lu
GraphMemix introduces a combinatorial‑optimization graph memory framework that organizes long‑term multimodal agent memory as query‑aware evidence forests. It constructs candidate graphs by expanding seed memories through schema and semantic relations, then decouples evidence utility from anchor‑conditioned relation verification to reduce redundancy, and finally optimizes a forest‑format memory context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.
GraphMemix introduces a combinatorial‑optimization graph memory framework that constructs query‑aware evidence forests for long‑term multimodal agent memory. It expands seed memories via schema and semantic relations, decouples memory support from relation verification to reduce redundancy, and optimizes a forest‑format context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.
By Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.
By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang
arXiv:2609.00551v1 Announce Type: cross
Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, s...
By Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng