arXiv:2607. 02781v1 Announce Type: cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates.
By Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum
Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.
By Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot, Amit Dhanda, Aman Chadha, Kapil Wanaskar, Vinija Jain, Amitava Das
arXiv:2505. 10892v2 Announce Type: replace Abstract: Post-training LLMs with RLHF and preference optimization methods (e.
By Akhil Agnihotri, Rahul Jain, Deepak Ramachandran, Zheng Wen
Mixture-Trained Merging (MTM) is a method for creating unified language models that combine multiple objectives—such as mathematics, code, instruction following, and controllable thinking—into a single parameter set. Instead of sequentially post‑training on each objective, MTM trains each branch on a mixture of objectives, ensuring that the branches remain compatible in weight space and can be merged without degrading performance. The approach iteratively refines merge coefficients using low‑cost evaluations and multi‑objective Bayesian optimization, outperforming naive merging and preserving distinct behaviors across domains.
By SeongHyeon Kim, Chaeyun Jang, Seungyoo Lee, Jiyeon Ham, Yunju Bak, Boseop Kim, Juho Lee
UniPolicy is a unified objective‑specific policy framework for search advertising that jointly optimizes relevance, click propensity, and commercial value. It uses objective‑aware prefix tokens, sparse MoE‑LoRA routing, and residual FFNs to decouple parameters within a shared backbone, and constructs pairwise preferences from multi‑stage behavioral feedback to strengthen clicked candidates. In large‑scale offline tests and a 7‑day online A/B test, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32% while keeping serving latency stable.
By Kun Yao, Yuhang Zhou, Yichi Zhang, Zeliang Tong, Shengri Xue, Haitao Wang, Siyu Lu, Qianlong Xie, Xingxing Wang
arXiv:2608. 16072v1 Announce Type: cross Abstract: Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners.
By Yixuan Wang, Yifei Chen, Haichao Zhang, Haozheng Luo, Xander Wu, Jie Ni, Yun Fu, Nuno Vasconcelos, Yijiang Li