arXiv Machine Learning

SHALA-LLM: Smartly Handling Ambiguous Labels in Aligning LLMs

arXiv:2606. 05376v1 Announce Type: new Abstract: Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challenging disagreements across human annotators.

arXiv Computation and Language
Sep 2

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani
arXiv Machine Learning
Jun 9

TinyJudge: Unverifiable Constraint Alignment via Lightweight Specialist Ensembles

arXiv:2606. 07520v1 Announce Type: cross Abstract: Instruction Following (IF) is a core capability of LLMs, requiring strict adherence to diverse constraints, ranging from verifiable ones (e.

By Yirong Zeng, Yufei Liu, Xiao Ding, Yutai Hou, Yuxian Wang, Wu Ning, Haonan Song, Dandan Tu, Qixun Zhang, Yuxiang He, Bibo Cai, Ting Liu
arXiv AI
Sep 25

Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures

The paper introduces Jev, a reinforcement‑learning‑trained model that provides calibrated probability answers to typed questions about a single input in one call. Jev is evaluated on RLCDAlignBench, a benchmark covering ten alignment failures across 44 tests and five target models, achieving a median AUROC of 0.886 zero‑shot and outperforming supervised baselines on most tasks. The study shows that question wording has little impact, while contextual fields that encode labels are more influential, and that Jev matches human‑label agreement while being 63× cheaper than LLM‑judge scorers.

By Ruoqi Guo, Yi Liu, Gelei Deng, Yuekang Li, Lida Zhao, Yutao Wu, Simin Chen, Ying Zhang, Leo Yu Zhang
Hugging Face Trending Papers
6d ago

Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks

The paper examines whether model uncertainty aligns with human disagreement on vision tasks. Using multi‑annotator datasets (FER+ and CIFAR‑10H), the authors find that pretrained models rarely reflect the ambiguity humans perceive, with weak correlations between model confidence and human disagreement. Predictive multiplicity offers only modest improvement, indicating that common uncertainty metrics fail to flag ambiguous cases.