arXiv AI By Zhisheng Chen, Tingyu Wu, Zijie Zhou, Zhengwei Xie, Jinhan Li, Ziyan Weng, Liang Lin, Jingwei Song, Zikai Xiao, Yingwei Zhang

PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Vision-Language Models

Read the original on arXiv AI →

arXiv:2602. 00415v2 Announce Type: replace Abstract: Memory is not merely a storage mechanism for intelligent systems, but a structure for organizing evidence and constraining belief.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 27

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that organizes long‑term multimodal agent memory as query‑aware evidence forests. It constructs candidate graphs by expanding seed memories through schema and semantic relations, then decouples evidence utility from anchor‑conditioned relation verification to reduce redundancy, and finally optimizes a forest‑format memory context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

arXiv AI
Aug 28

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that constructs query‑aware evidence forests for long‑term multimodal agent memory. It expands seed memories via schema and semantic relations, decouples memory support from relation verification to reduce redundancy, and optimizes a forest‑format context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

By Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
Hugging Face Trending Papers
Jun 9

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing memory paradigms represent each memory item in raw text and image forms, so retrieval-based systems must pass the retrieved text or images to the generation LLMs/VLMs, resulting in high token consumption and storage pressure, making it unaffordable for resource-constrained applications.