arXiv AI

PolarMem: A Training-Free Polarized Latent Graph Memory for Verifiable Vision-Language Models

arXiv:2602. 00415v2 Announce Type: replace Abstract: Memory is not merely a storage mechanism for intelligent systems, but a structure for organizing evidence and constraining belief.

Hugging Face Trending Papers
Aug 27

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that organizes long‑term multimodal agent memory as query‑aware evidence forests. It constructs candidate graphs by expanding seed memories through schema and semantic relations, then decouples evidence utility from anchor‑conditioned relation verification to reduce redundancy, and finally optimizes a forest‑format memory context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

arXiv AI
Aug 28

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

GraphMemix introduces a combinatorial‑optimization graph memory framework that constructs query‑aware evidence forests for long‑term multimodal agent memory. It expands seed memories via schema and semantic relations, decouples memory support from relation verification to reduce redundancy, and optimizes a forest‑format context within a maximum evidence budget. Experiments on four benchmarks show significant accuracy gains and a new Pareto frontier between accuracy and lifecycle cost.

By Geng Li, Yuhao Wang, Dong Li, Jianye Hao, Yuxin Peng
Hugging Face Trending Papers
Jun 9

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing memory paradigms represent each memory item in raw text and image forms, so retrieval-based systems must pass the retrieved text or images to the generation LLMs/VLMs, resulting in high token consumption and storage pressure, making it unaffordable for resource-constrained applications.

arXiv AI
Aug 7

ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion

arXiv:2608. 05833v1 Announce Type: new Abstract: Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimodal knowledge graph completion (MMKGC), where entities are associated with multiple modalities such as text and images.

By Jiafan Li, Mengxue Yang, Jiaqi Zhu, Liang Chang, Ying Li, Hongan Wang
arXiv AI
6d ago

Can Linguistic Reasoning Vectors Enhance Multimodal Reasoning Ability?

The paper introduces LIFT, a lightweight vector‑intervention technique that transfers reasoning capability from a base large language model (LLM) to a vision‑language model (VLM) without retraining the VLM backbone. LIFT defines Reasoning Vectors as differences in hidden states between a reasoning path with an explicit trace and a solver path without it, and injects these vectors into the VLM’s language‑side activations. Experiments on two VLMs across six reasoning benchmarks show that vectors derived from the base LLM consistently outperform those derived from the aligned VLM, indicating that the base LLM is a more effective source for recovering degraded reasoning. "whyItMatters":"The study demonstrates that a simple, frozen‑backbone intervention can partially restore reasoning abilities in multimodal models, highlighting the value of leveraging the original language model’s reasoning power."

By Ziyi Wang, Li Li, Aolin Zhou, Yankun Shen, Chonghan Liu, Shuxia Lin, Xu Yang
arXiv AI
Jul 21

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering

arXiv:2510. 01483v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly: because VLMs maintain no persistent memory or explicit spatial representation, all sampled frames must be re-processed for every new query.

By Mohamad Al Mdfaa, Svetlana Lukina, Timur Akhtyamov, Arthur Nigmatzyanov, Dmitrii Nalberskii, Sergey Zagoruyko, Gonzalo Ferrer