arXiv AI

Toward an Unbiased Collective Memory for Efficient LLM-Based Agentic 6G Cross-Domain Management

arXiv Machine Learning
Sep 1

A-MADiff: Attention-Guided Multi-Agent DRL with Diffusion Policies for Memory-Aware Task Orchestration in Mobile AIGC Networks

arXiv:2608.29255v1 Announce Type: cross Abstract: Artificial Intelligence-Generated Content (AIGC) services employ Generative AI (GenAI) models to automatically generate diverse content. Mobile AIGC...

By Chongzhi Wu, Zhengtao Li, Jiawen Kang, Jinbo Wen, Xiaohuan Li, Maomao Zhang, Ekram Hossain
arXiv Machine Learning
Jul 10

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents

arXiv:2607. 08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and storing what happened in previous turns and sessions.

By Ashwin Gerard Colaco, Nada Lahjouji
arXiv AI
6d ago

Endogenous Information in Routing Games: Memory-Constrained Equilibria, Recall Braess Paradoxes, and Memory Design

The paper investigates routing games where travelers choose routes based on remembered or surfaced alternatives rather than a fixed set of actions. It introduces a tractable design theory for endogenous recall, linking a finite‑memory micro model—where each traveler updates a memory state via a logit rule and policies like LRU—to a stationary salience model that assigns route‑specific weights. The authors prove existence and uniqueness of a Forgetful Wardrop Equilibrium, develop algorithms for network design under budget constraints, and identify a Recall Braess Paradox where better recall can worsen equilibrium delay.

By Saad Alqithami
arXiv AI
Sep 17

Where Should Agents Live? Energy-Memory Characterization of Agentic AI for the Edge-Cloud Continuum

The paper introduces agentic-eCAL, an extension of the Energy Cost of AI Lifecycle metric to evaluate multi‑agent AI workflows across the edge‑cloud continuum. By combining a two‑rate energy model with OSI‑layer transport analysis, the authors quantify that inter‑agent text transfer accounts for only 0.25% of total workflow energy, highlighting that the main energy cost lies in additional inference and context processing triggered by communication. The study uses extensive GPU benchmarks on NVIDIA A100/H100 with 16 open‑weight models and 8 orchestration topologies to validate the metric and explore placement implications.

By Carolina Fortuna, Vid Han\v{z}el, Tim Strnad, Bla\v{z} Bertalani\v{c}