arXiv Machine Learning

ExpBoN: Exponential-Noise Best-of-$n$ for Efficient Test-Time LLM Alignment

arXiv AI
Sep 2

Bandits in Prod: Hyperparameter Optimization at Inference Time

The paper introduces Online Hyperparameter Optimization (OHPO), framing it as an infinitely many‑armed bandit problem over mixed and conditional search spaces. It proposes the IMABO framework, which couples any bandit policy with any oracle for proposing new configurations, and presents IMOSS—a restart‑free anytime policy with provable regret bounds. Experiments show that IMABO, combined with practical oracles such as TPE, an incumbent‑mutation oracle, and a pretrained tabular foundation model, outperforms random search across a range of settings from classical ML models to LLM‑based agents.

By Louis Abraham, Tuan-Anh Nguyen, Nicolas Devatine
arXiv Machine Learning
Sep 22

Swiss-Knife: A Framework for Reconfigurable Externalised Multi-Objective Alignment at Decode Time

Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.

By Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot, Amit Dhanda, Aman Chadha, Kapil Wanaskar, Vinija Jain, Amitava Das
arXiv Machine Learning
Aug 28

GRAS: Guided Reduced-Variance Proposals and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

The paper introduces GRAS, a method that improves training‑free reward alignment for discrete diffusion models by reducing variance in guided proposals and adapting the resampling temperature during search. It achieves this without adding denoiser cost, using Rao‑Blackwellized estimates for differentiable rewards and a leave‑one‑out baseline for non‑differentiable ones. Experiments on regulatory DNA and protein design show GRAS outperforms existing training‑free techniques and rivals reward‑fine‑tuned models.

By Kwanyoung Kim
arXiv Machine Learning
Aug 28

Privacy Without Regret: Differentially Private Inference-Time Alignment

The paper introduces Private Best-of-N (PrivBoN), a method that adds calibrated Gumbel noise to reward scores during inference-time alignment, achieving both ε-differential privacy and KL-regularized alignment. When the privacy budget exceeds a critical threshold ε*, the noise becomes regret-optimal, matching the theoretical alignment skyline. The authors also propose Private Inference-Time Pessimism (PrivITP), which uses χ^2-regularized rejection sampling and a two-phase Gaussian mechanism to provide ex-post (ε,δ)-DP with a privacy cost independent of the number of responses, and demonstrate that both methods outperform standard Best-of-N across multiple models and datasets.

By Ishi Jain, Nandini Bhattad, Sayak Ray Chowdhury
arXiv Machine Learning
Aug 20

Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning

The paper introduces the concept of decision‑metric alignment, which ensures that Euclidean distance to a goal latent in JEPA‑style latent world models correctly ranks action sequences for model‑predictive control. It proposes two metrics—Plan‑Real Spearman and CEM‑stage Spearman—to evaluate latent–real rank agreement, and identifies encoder distortion, terminal rollout error, and candidate margins as key factors affecting alignment. Building on these insights, the authors present DA‑LeWM, an enhanced latent world model that incorporates inverse‑dynamics and demonstration‑conditioned goal‑action heads, leading to faster convergence and higher online success rates compared to the baseline LeWM while maintaining similar probe scores.

By Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
arXiv AI
3d ago

Multi-LLM Collaborative Alignment via Stackelberg Games

The paper introduces Stackelberg Alignment, a leader‑follower framework that lets a pool of language models collaborate and improve by learning from each other’s responses. An EXP3 bandit leader adaptively selects instructions based on difficulty and discriminability, while the models act as followers, evaluating peers and learning via DPO or GRPO with Elo‑style reputation weighting and opponent matching. Experiments on diverse benchmarks show that this adaptive curriculum outperforms static baselines by up to 12‑25% and improves multi‑LLM evolution.

By Christina Hahn, Shangbin Feng, Dean Light, Swastik Roy, Hila Gonen, Yulia Tsvetkov