arXiv AI

Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference

arXiv:2608. 06085v1 Announce Type: new Abstract: Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random.

arXiv AI
Aug 21

Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

arXiv:2608. 19760v1 Announce Type: cross Abstract: Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- identifies which steps causally matter better than chance.

By Haiyue Zhang
arXiv AI
2d ago

Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model

The paper investigates whether confidence scores from a black-box decision model, Jev, truly reflect missing knowledge. Using over 15 public datasets and 6 synthetic task families, the authors find that while Jev’s confidence is calibrated on familiar closed-choice tasks, it fails to indicate when the model lacks relevant information—assigning high confidence to salient options even without answer-relevant data and overestimating accuracy on news beyond its knowledge boundary. Targeted yes/no questions about whether an outcome is settled or whether evidence suffices provide sharper indicators of knowledge gaps, but only when surface cues are controlled.

By Sharath M Shankaranarayana, Davor Runje, Jan Jannink
arXiv Machine Learning
Jun 30

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

arXiv:2606. 28963v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated.

By Eun Cheol Choi, Youngrae Kim, Prabhu Pugalenthi, Hong-En Chen, Bo-Ruei Huang