arXiv AI

TokenPilot: Cache-Efficient Context Management for LLM Agents

arXiv:2606. 17016v1 Announce Type: cross Abstract: As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs.