arXiv AI

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

arXiv:2510. 17426v3 Announce Type: replace-cross Abstract: The "alignment tax" of post-training is typically framed as a drop in task accuracy.

arXiv Computer Vision
Sep 11

Task Alignment: A Simple Proxy for Practical Model Merging Across Diverse Vision Tasks

The paper introduces the task alignment proxy, a method that accelerates hyperparameter selection for merging models fine‑tuned on diverse vision tasks. It addresses the challenge of training heterogeneous decoders, which makes traditional downstream performance evaluation costly. By using the proxy, the authors demonstrate that model merging can be applied efficiently to multi‑task vision models beyond CLIP‑based classification.

By Pau de Jorge, C\'esar Roberto de Souza, Bj\"orn Michele, Mert B\"ulent Sar{\i}y{\i}ld{\i}z, Philippe Weinzaepfel, Florent Perronnin, Diane Larlus, Yannis Kalantidis
arXiv AI
Jul 7

Multi-Way Representation Alignment

arXiv:2602. 06205v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.

By Akshit Achara, Tatiana Gaintseva, Mateo Mahaut, Pritish Chakraborty, Viktor Stenby Johansson, Melih Barsbey, Emanuele Rodol\`a, Donato Crisostomi
arXiv AI
4d ago

Alignment Forecasting: Predicting Misalignment From Training Data

The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.

By Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
arXiv Computation and Language
Sep 1

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.

By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang
arXiv AI
Jul 17

Decoupled Alignment for Robust Plug-and-Play Adaptation

arXiv:2406. 01514v4 Announce Type: replace-cross Abstract: We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tuning or reinforcement learning from human feedback.

By Haozheng Luo, Jiahao Yu, Wenxin Zhang, Jialong Li, Chenghao Qiu, Yimin Wang, Eric Hanchen Jiang, Jerry Yao-Chieh Hu, Yan Chen, Binghui Wang, Xinyu Xing, Han Liu