Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no soundness statement, and certified approaches budget request-...
arXiv:2608. 05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression.
By Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
arXiv:2609.23886v1 Announce Type: new
Abstract: Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears wit...
By Zehua Cheng, Wei Dai, Jiahao Sun
ServeGuard is a supply‑chain primitive that allows a publisher to ship a proof‑carrying adapter for an open‑weight language model, proving in zero‑knowledge that the adapter contains no hidden backdoor channel in the monitor’s blind subspace. The proof is inexpensive because it relies on a deterministic function of the public base model, and the served residual is the model’s own public floor. The system lets consumers or regulators verify the absence of this class of hidden channels without revealing the certified read factor or trusting the publisher.
By Dominik Dahlem, Rui Vieira
arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
arXiv:2609.07162v1 Announce Type: new
Abstract: Several properties safety monitors are asked to certify, among them cross-tenant noninterference, sandbagging and evaluation awareness, are 2-safety hy...
By Xin Xu
The paper argues that large language model (LLM) providers, constrained by compute, often degrade service during congestion by routing queries to smaller models, cutting reasoning effort, or truncating context. It shows that this practice misrepresents costs because degraded answers can fail, leading to retries that inflate traffic or churn that erodes lifetime value. By modeling inference allocation with newsvendor, retry, and queueing frameworks, the authors derive a ‘shadow price of intelligence’ that quantifies the marginal value of each query, revealing that throttling under congestion acts as a demand lever rather than a cost lever.
By Elioth Sanabria
arXiv:2606. 07316v2 Announce Type: replace-cross Abstract: Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label?
By Haoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang
PAGE is a partition‑aware gated KV‑cache eviction method that reframes eviction as a per‑input admission decision. It uses a single label‑free scalar— the early‑to‑late drop in pairwise top‑k head agreement—to classify inputs into a capacity‑bound class (where eviction is catastrophic) and a dilution‑prone class (where eviction is safe or beneficial). By thresholding this drop, PAGE applies a base evictor only when necessary, reducing the harm rate in the capacity‑bound regime from 0.75 to 0.026 and achieving a 29× improvement across four models and benchmarks without retraining the evictor.
By Pankaj Kumar, Subhankar Mishra
The paper shows that Worst‑Case Optimal Recovery (OR) and Bayesian learning solve the same Gaussian‑quadratic‑Hilbert problems, linking the radius of information to a nugget‑optimized Gaussian process posterior variance. It evaluates three Bayesian systems, demonstrating that OR can outperform Bayesian methods in certain calibration and reproducibility metrics, yet split‑conformal and other approaches can beat OR in interval scoring, especially under covariate shift. The authors propose matching the guarantee tool to the data regime and auditing that regime first.
By Gordei Verbii
arXiv:2609.25686v1 Announce Type: cross
Abstract: Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent...
By Chenyu Zhang, Wonbin Kweon, Jiawei Han
The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.
By Vishnu Bindu Balachandran