arXiv AI By Xueshu Chen, Yan Wang, Zihao Xue, Jiefu Li, Zhenfang Liu, Jayden Chen, Zhen Bi, Jungang Lou

C3M: Cross-Session Multimodal Memory Maintenance for Long-Horizon Tasks

Read the original on arXiv AI →

C3M is a cross‑session multimodal memory system designed for long‑horizon tasks that must preserve and retrieve evidence across sessions within a limited, query‑blind memory budget. It maintains a bounded active index of source text‑image evidence, using relation‑aware updates to keep safe redundancy while preserving complementary and incompatible records. At query time, budgeted routing selects useful index pages and expands their associated source evidence under a fixed reader budget, creating a compact, provenance‑preserving memory that retains temporal distinctions and source links for reliable downstream reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 27

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that organizes long‑term multimodal agent memory as query‑aware evidence forests. It constructs candidate graphs by expanding seed memories through schema and semantic relations, then decouples evidence utility from anchor‑conditioned relation verification to reduce redundancy, and finally optimizes a forest‑format memory context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

arXiv AI
Aug 28

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that constructs query‑aware evidence forests for long‑term multimodal agent memory. It expands seed memories via schema and semantic relations, decouples memory support from relation verification to reduce redundancy, and optimizes a forest‑format context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

By Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
arXiv Computer Vision
Aug 25

FOVEA: Focused On-Demand Visual Evidence Adaptation for Cache-Friendly Multimodal Speculative Decoding

FOVEA introduces a cache‑friendly, on‑demand visual evidence adaptation for multimodal speculative decoding, enabling a lightweight draft model to dynamically retrieve a bounded subset of visual memory based on a cumulative‑mass rule. The retrieved visual readout is fused with the draft hidden state via a lightweight gated residual correction, avoiding the insertion of visual tokens into the autoregressive context. Experiments on various vision‑language backbones and benchmarks show that FOVEA improves draft acceptance and speeds up end‑to‑end decoding by up to 2.13× compared to traditional autoregressive decoding.

By Hengjie Zhu, Dayan Wu, Zihao Zhang, Xinze Liu, Jingxuan Yu, Peng Fu, Zheng Lin, Weiping Wang, Ding Wang
arXiv AI
Sep 2

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

arXiv:2609.00551v1 Announce Type: cross Abstract: Multimodal memory offers a scalable interface for long-video question answering, but existing methods often retrieve captions, frames, transcripts, s...

By Yijun Chen, Yaqi Zheng, Yanya Li, Boyi Xiao, Buqiang Xu, Shuofei Qiao, Jizhan Fang, Xinle Deng, Yunzhi Yao, Xuehai Wang, Liuxin Zhang, Hui Li, Huajun Chen, Shumin Deng