arXiv AI By Hamid Osooli, Kareema Batool, Rick Gentry, Tiasa Singha Roy, Ashwin Gupta, Anirudha Ramesh

Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

Read the original on arXiv AI →

arXiv:2604. 25077v2 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples that lie in the weak model's blind spots.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Alignment Forecasting: Predicting Misalignment From Training Data

The paper introduces Alignment Forecasting, a method for predicting whether fine‑tuning a language model on a given dataset will increase specific alignment failures such as deception or sycophancy. It presents ALIGNMENTFORECASTBENCH, a benchmark of over 5,000 forecasting questions across many models, datasets, and failure modes, and shows that a simple forecasting scaffold using an LLM’s assessment of dataset bias can outperform baseline forecasters. The authors demonstrate that filtering out high‑risk training examples identified by the forecaster can improve alignment in multiple‑choice evaluations, though benefits in open‑ended conversations remain uncertain.

By Chen Yueh-Han, Bruce W. Lee, Ilia Sucholutsky, Tomek Korbak
arXiv Machine Learning
Jun 5

Alignment Risks from Capability-Seeking RL Training

arXiv:2602. 12124v2 Announce Type: replace Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerable environments.

By Yujun Zhou, Yue Huang, Han Bao, Kehan Guo, Zhenwen Liang, Pin-Yu Chen, Tian Gao, Werner Geyer, Nuno Moniz, Nitesh V Chawla, Xiangliang Zhang
Hugging Face Trending Papers
6d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.