arXiv AI By Bruce Changlong Xu, Adarsh Kumarappan, Mu Zhou

Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation

Read the original on arXiv AI →

arXiv:2606. 09864v1 Announce Type: cross Abstract: Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely focus on measuring perplexity and accuracy without assessing the safety impact.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.