arXiv AI

Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents

arXiv:2606. 10209v1 Announce Type: new Abstract: Large language models deployed as autonomous agents for enterprise workflows face a key challenge: verbose tool responses from enterprise systems can cause context overflow, stale-state errors, and high inference cost.

arXiv AI
Aug 19

Token Optimization and Context Window Management in Multi-Agent AI Workflows

The paper "Token Optimization and Context Window Management in Multi‑Agent AI Workflows" introduces a practitioner framework that reduces token usage and latency in multi‑agent AI systems. It outlines six patterns—context stratification, fetch‑once/process‑locally architecture, schema‑contracted prompts, token‑aware fallback chains, semantic caching, and inter‑agent communication compression—and reports a 60‑70% token reduction and a 61‑116 second cold‑load latency improvement in production. A controlled study on relevance‑contrast context shows that mixing high‑ and low‑relevance items in prompts can improve relevance accuracy by up to +0.084. whyItMatters":"The work provides concrete, repeatable engineering patterns that bridge research and production, enabling faster, cheaper, and more reliable AI workflows."

By Dvir Shamay
Hugging Face Trending Papers
Jul 23

Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems

Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations.

arXiv AI
Sep 16

Protocol-Preserving Context Trimming for Agentic Workflows: Benefits, Failure Regimes, and Budget Guardrails

The paper evaluates five context‑trimming strategies for agentic large language model workflows, comparing them on metrics such as task success, protocol adherence, token savings, and latency. Conventional trimming methods save about 60% of tokens but achieve lower success rates, while protocol‑aware trimming raises success to 92.2% and adaptive guardrails further improve it to 96% success with 56% token savings. The study shows that preserving protocol‑critical state is more important than aggressive token removal, and that adaptive guardrails enhance efficiency, scalability, and reliability for long‑horizon agentic systems.

By Harish Gaggar
arXiv AI
Aug 6

ContextWeave: A Real-World Workflow Benchmark

arXiv:2608. 04830v1 Announce Type: new Abstract: Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering.

By Bo Wang, Yuqian Yao, Enxi Wang, Luozhijie Jin, Yang Liu, Yiran Suo, Yuxuan Cai, Enyu Zhou, Yufei Gao, Honglin Guo, Tianyu Huai, Li Ji, Zhikai Lei, Bufan Li, Lizhi Lin, Jinxiu Liu, Jie Yang, Jiazheng Zhou, Maosen Zhou, Pengfang Qian, Shichun Liu, Guanshan Liu, Hao Zheng, Yunhao Yu, Hang Yan, Jihua Kang, Xinchi Chen, Xipeng Qiu
arXiv Computation and Language
Aug 31

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot is a proactive context‑management framework designed to improve long‑horizon agentic reasoning with large language models. It expands the toolset to include planning, long‑term memory, and soft context offloading, and introduces a reinforcement‑learning strategy that focuses on critical editing decisions and assigns action‑level advantages. Experiments on long‑context QA and deep search tasks demonstrate that ContextPilot achieves stronger performance with a more compact working context, outperforming existing baselines across various base models and benchmarks.

By Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun
Hugging Face Trending Papers
Aug 5

ContextWeave: A Real-World Workflow Benchmark

Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We introduce ContextWeave, a longitudinal benchmark that evaluates whether recalled experience improves downstream agent performance in realistic office-work streams.

arXiv AI
Sep 17

RideWay: Benchmarking Efficient Task Completion for Tool-Using Language Agents

RideWay is a new benchmark that evaluates ride‑hailing language agents not just on task completion but on interaction efficiency. It introduces the Efficiency Utility metric, which penalizes agents for excessive tool calls and user‑facing turns relative to a task‑specific reference effort, with human preferences used to calibrate the penalties. Across 58 tasks and 24 models, the metric shows that extra dialogue is penalized more heavily than extra tool use, and it achieves high accuracy in distinguishing trajectories that differ in turns but struggles when differences are only in tool calls.

By Qingnuan Han, Boli Fang, Mingzhi Hou, Claire Liu