arXiv AI By Fanzhe Wei, Li Liu

WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization

Read the original on arXiv AI →

arXiv:2607. 28699v1 Announce Type: cross Abstract: KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the request it is serving right now.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 31

A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint

The paper introduces a method for allocating bit-widths to individual layers in Gemma-3-1B to maximize performance (latency reduction) while staying within a specified quality-degradation budget. Using a layer sensitivity profile from SA-PTQ and TensorRT-LLM, the authors evaluate 13 W8A8 variants on an RTX 5090, finding that FFN 5+5 with lm_head yields an 11.0% latency reduction with negligible quality loss. They also discuss trade-offs for Attention layers and propose further optimizations such as fused INT8 attention kernels and FP8 usage.

By Artem Safronov