arXiv Machine Learning

Closed-Loop Bayesian Bandit Encoder with GRAND Receiver for a Bursty Interference Channel

arXiv:2607. 15404v1 Announce Type: cross Abstract: Interleaving mitigates burst errors but introduces decoding delay and removes temporal error structure that a channel-aware decoder could exploit.

arXiv Machine Learning
1d ago

Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift

The paper studies how an eavesdropper can adaptively attack quantum key distribution (QKD) systems when channel noise and device drift vary over time. By modeling the attack as a constrained Markov decision process and using reinforcement learning to jointly search gate structures and rotation angles, the authors construct compact attack circuits that perform near the theoretical upper bound for both device‑independent E91 and BB84 protocols under realistic noise models. The results show that adaptive attacks can significantly increase the eavesdropper’s information compared to fixed‑circuit strategies, and that the learned attacks recover known optimal cloners and key‑rate bounds.

By Marcel Mordarski, Benjamin Gras, Abdelrahman Shehata, Daniel Budina, Roberto Bondesan
arXiv AI
Aug 26

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

The paper introduces a Bayesian self‑escalation strategy for hierarchical large‑language‑model agents, allowing an agent to detect during its own reasoning that it is unlikely to succeed and hand control over to a stronger model. The authors formalise this as an optimal‑stopping problem over a learned competence posterior, derive a myopic escalation threshold, and prove that the optimal policy is a time‑varying threshold without assumptions on the raw signal. They provide theoretical guarantees—including a 1/√n regret decay with n calibration trajectories—and validate the approach in simulations and a real‑model code‑generation cascade, showing that the escalation frontier outperforms post‑hoc routing at equal cost. whyItMatters":"The study offers a principled, theoretically grounded method for agents to dynamically decide when to seek stronger models, potentially improving efficiency and reliability in hierarchical LLM systems."

By Nadeem Shaikh
arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv AI
Jul 14

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

arXiv:2607. 11317v1 Announce Type: new Abstract: Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought.

By El Hassane Ettifouri (Novelis Research, Paris, France), Ayoub Belfatmi (Novelis Research, Paris, France), Mahaman Sanoussi Yahaya Alassan (Novelis Research, Paris, France), Walid Dahhane (Novelis Research, Paris, France)