arXiv:2605. 07914v2 Announce Type: replace Abstract: Sharpness-aware and gradient-alignment methods have been shown to improve generalization, however each family of methods targets a single geometric property of the loss landscape, while ignoring the other.
By Aristotelis Ballas, Christos Diou
The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.
By Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
arXiv:2511. 16992v3 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human values often involves balancing multiple, conflicting objectives such as helpfulness and harmlessness.
By Fatemeh Nourzad, Amirhossein Roknilamouki, Eylem Ekici, Jia Liu, Ness Shroff
arXiv:2607. 02781v1 Announce Type: cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates.
By Yaswanth Chittepu, Ativ Joshi, Sohini Chintala, Scott Niekum
arXiv:2609.21899v1 Announce Type: new
Abstract: Best-of-$n$ (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade...
By Yanxiao Liu, Sicheng Wan, Deniz G\"und\"uz
arXiv:2606. 18650v1 Announce Type: new Abstract: As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories.
By Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia Jiang
The paper introduces the concept of decision‑metric alignment, which ensures that Euclidean distance to a goal latent in JEPA‑style latent world models correctly ranks action sequences for model‑predictive control. It proposes two metrics—Plan‑Real Spearman and CEM‑stage Spearman—to evaluate latent–real rank agreement, and identifies encoder distortion, terminal rollout error, and candidate margins as key factors affecting alignment. Building on these insights, the authors present DA‑LeWM, an enhanced latent world model that incorporates inverse‑dynamics and demonstration‑conditioned goal‑action heads, leading to faster convergence and higher online success rates compared to the baseline LeWM while maintaining similar probe scores.
By Jiawei Wang, Ke Rui, Yushen Zuo, Yichun Feng, Minglei Li
The paper investigates how latent world models (specifically JEPA-style models) use Euclidean distance to a goal latent as a cost for model‑predictive control (MPC). It introduces two metrics—Plan‑Real Spearman and CEM‑stage Spearman—to evaluate how well latent‑space distances align with real‑task progress, a property termed decision‑metric alignment. By identifying encoder distortion, terminal rollout error, and candidate margins as key factors, the authors propose DA‑LeWM, which augments the base model with inverse‑dynamics and demonstration‑conditioned goal‑action heads, leading to faster convergence and higher online success while maintaining similar probe scores.
arXiv:2605. 24395v2 Announce Type: replace Abstract: Alignment plays a fundamental role in many machine learning problems, such as multi-network analysis, multimodal learning, and point cloud registration.
By Qi Yu, Ruizhong Qiu, Zhichen Zeng, My T. Thai, Huan Liu, Hanghang Tong
The paper studies algorithms for computing the Entropic Gromov-Wasserstein (EGW) distance, a measure of discrepancy between metric measure spaces. It introduces Averaged Mirror Descent (AMD), which averages successive Mirror Descent steps and is proven to converge for any cost function, and shows that a dual gradient method with a fixed step size also converges for arbitrary costs, even when iterations are inexact. Empirical comparisons demonstrate that both AMD and the dual gradient method succeed on cases where classical Mirror Descent fails.
By Joanna Marks, Gabriel Rioux, Riccardo Passeggeri
arXiv:2609.06893v1 Announce Type: cross
Abstract: Direct preference optimization DPO is a promising offline approach for aligning large language models (LLMs) due to its simplicity, computational eff...
By Wenbo Zhang, Wenzhuo Zhou, Hengrui Cai, Zhengling Qi
arXiv:2310. 15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients.
By Zhen Qin, Zhishuai Liu, Pan Xu