arXiv:2606. 31023v1 Announce Type: cross Abstract: Hard-constrained sequential decision systems have no certified way to spend the test-time compute of modern AI: executing the multi-step drafts of a learned policy or a frozen LLM forfeits the feasibility guarantee a trusted solver provides, while invoking the solver at every step forfeits the speed the AI offers.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
ServeGuard is a supply‑chain primitive that allows a publisher to ship a proof‑carrying adapter for an open‑weight language model, proving in zero‑knowledge that the adapter contains no hidden backdoor channel in the monitor’s blind subspace. The proof is inexpensive because it relies on a deterministic function of the public base model, and the served residual is the model’s own public floor. The system lets consumers or regulators verify the absence of this class of hidden channels without revealing the certified read factor or trusting the publisher.
By Dominik Dahlem, Rui Vieira
arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.
By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.
By Amit Singh Bhatti, Vishal Vaddina
arXiv:2606. 06460v3 Announce Type: replace-cross Abstract: Autonomous LLM agents increasingly hold real credentials and operate infrastructure with no human in the loop, yet operators have no standard way to tell an agent a resource is off-limits, or to ask a running agent to stand down: access controls either admit it or hard-fail it.
By Thamilvendhan Munirathinam
arXiv:2609.27822v1 Announce Type: cross
Abstract: A common multi-agent design asks agents to report confidence and lets the highest-scoring agent speak next, implicitly using one scalar both to route...
By Jingyan Jiang, Huihuo Zheng, Rajeev Thakur, Chih-Hsuan Yang