arXiv AI By Rui Zhu, Weiheng Bai, Qiushi Wu, Yang Ren, Haixu Tang, Yuchu Liu

How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment

Read the original on arXiv AI →

The paper addresses the memory bottleneck in reinforcement learning for large language models caused by the large Key-Value (KV) cache during rollout phases. It highlights that while KV cache compression can reduce memory usage, it introduces a significant off‑policy bias that standard statistical corrections cannot adequately mitigate. The authors argue that even tiny compression errors are amplified by RL’s instability, leading to inefficient learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.