arXiv AI

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

The paper investigates item-sensitivity—whether a language model’s choice depends on the specific input—in a forced-choice signalling task derived from the board game Deception: Murder in Hong Kong. Across seven models, two families, a post‑training ablation, and three scoring rules, every tested cell shows item‑sensitivity, yet many are statistically indistinguishable from random choice and some perform worse than random. The authors term this phenomenon "consistency without alignment" and argue it undermines evaluations that rely solely on item‑sensitivity, permutation consistency, or self‑consistency without an independent reference.

arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi
arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur