← Back to all news
arXiv Machine Learning October 7, 2026 By Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim

PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • agents
  • fine-tuning
  • efficiency
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 2

LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents

arXiv:2602. 01053v2 Announce Type: replace Abstract: Role specialization in multi-LLM agent systems is often realized via multi-LoRA, where agents share a pretrained backbone and differ only by lightweight adapters.

By Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim
llmsagentsfine-tuningefficiencybenchmarks
More like this →
arXiv Machine Learning
1d ago

KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems

arXiv:2609.34060v2 Announce Type: replace Abstract: Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However...

By Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang, Jae-Joon Kim
agentsefficiencymultimodal
More like this →
arXiv AI
Aug 18

Learning Agent Execution for KV-Cache Management in Agentic Serving

arXiv:2608. 14624v1 Announce Type: new Abstract: Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequence of specialized agents.

By Rui Zhang, Chaeeun Kim, Shaoting Feng, Kuntai Du, Yuhan Liu, Yi Zhong, Cheng-Wei Ching, Junchen Jiang, Liting Hu
llmsagentsefficiency
More like this →
arXiv AI
Jun 6

RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

arXiv:2606. 06256v1 Announce Type: new Abstract: As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastructure.

By Yang Liu, ZhaoKai Luo, HuaYi Jin, ZhiYong Wang, RuoZhou He, BoYu Wang, Guanjie Chen, Junhao Hu
llmsfine-tuningefficiency
More like this →
arXiv AI
Oct 1

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

arXiv:2609.32259v2 Announce Type: replace Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires ea...

By Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy
llmsagentsnlpefficiencybenchmarks
More like this →
arXiv Computation and Language
Aug 21

ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM Agents

arXiv:2608. 19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states.

By Yichu Fang, Sitong Wei, Haozhe Hu, Xiaoyu Shen
llmsagentsnlpefficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea