arXiv Machine Learning

A Probabilistic Approach for Model Alignment with Human Comparisons

The paper proposes a two‑stage framework, SL+LHF, that first learns low‑dimensional representations from noisy labeled data and then refines model alignment using human comparison feedback via a probabilistic bisection approach. It introduces the label‑noise‑to‑comparison‑accuracy (LNCA) ratio to theoretically identify when this framework outperforms pure supervised learning, showing that trading labels for comparisons reduces sample complexity when labels are scarce. Experiments on a high‑dimensional crowdfunding prediction task and an Amazon Mechanical Turk study confirm that incorporating human or large language model evaluators improves accuracy under a fixed query budget.

arXiv Computation and Language
Sep 2

Post-hoc Alignment of LLM-judges to Human Judgment Distribution

The paper introduces NAPHA, a lightweight post‑hoc alignment method that improves large language model (LLM) predictions of human judgment distributions (HJD) by matching LLM output distributions to HJD through entropy‑based class assignment and specialized alignment models. Experiments on five datasets show that while LLMs perform near human‑level on hard‑label tasks, they struggle with soft‑label predictions, and NAPHA consistently enhances soft‑label accuracy, especially on high‑entropy instances. The study also demonstrates that better entropy class prediction can further boost NAPHA’s effectiveness.

By Sebastian Steindl, Nikos Voskarides, Alberto Gasparin, Diego Marcheggiani
arXiv Machine Learning
Jul 15

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

arXiv:2602. 07008v3 Announce Type: replace-cross Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence.

By Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Sanyi Zhang, Kangwei Liu, Shiming Liu, Zhangcheng Wang, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao
arXiv AI
2d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi