arXiv AI By Cris Huynh

Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random

Read the original on arXiv AI →

The paper investigates item-sensitivity—whether a language model’s choice depends on the specific input—in a forced-choice signalling task derived from the board game Deception: Murder in Hong Kong. Across seven models, two families, a post‑training ablation, and three scoring rules, every tested cell shows item‑sensitivity, yet many are statistically indistinguishable from random choice and some perform worse than random. The authors term this phenomenon "consistency without alignment" and argue it undermines evaluations that rely solely on item‑sensitivity, permutation consistency, or self‑consistency without an independent reference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi