arXiv AI

How much of a measured AI preference is the model, and how much is the instrument?

The paper investigates how much of an AI model’s expressed preferences are due to the model itself versus the instrument (prompt) used to elicit those preferences. By fixing the set of outcomes and models while varying five different prompting instruments across 15 welfare-related outcomes, the study finds that the ranking of outcomes is only moderately generalizable (coefficient 0.348) and that a single instrument’s preference provides little insight into another instrument’s results. The analysis shows that even after removing any single instrument, model, or a subset of outcomes, the overall preference estimate remains robust, yet the variability across instruments remains significant.

Hugging Face Trending Papers
Aug 18

Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

The paper investigates whether reader-specific differences in retrieval‑augmented generation (RAG) reflect reusable structure or merely input‑local interactions. By fixing query, evidence, task, scoring, and intervention, the authors find that nine readers disagree on the effect sign in 33% of cases, with reader×query interactions explaining 29.8% of utility variance. They further decompose heterogeneity into evidence activity, ordinal preference, and conditional signed direction, discovering that ordinal reader geometry is stable across multiple settings while signed geometry is task‑bounded, yet stable ordinal similarity does not predict cross‑reader intervention transfer.

arXiv Machine Learning
Aug 14

A Probe Direction Is a Property of Its Prompt

arXiv:2608. 13329v1 Announce Type: new Abstract: A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations.

By Valentin No\"el
arXiv AI
Jun 10

Does Capability Transfer to Subjective Behavior -- and Would Our Instruments Tell Us? A Self-Evolving, Trust-by-Construction Evaluation Paradigm

arXiv:2605. 27914v2 Announce Type: replace-cross Abstract: Benchmarking is mature where answers are verifiable -- math, code, reasoning -- but the fastest-growing uses of LLMs are subjective and human-facing: companionship, emotional support, counseling.

By Yuming (Rapheal), Huang, Yao Liu, Pengjie Ding, Lei Wang, Junchen Wan
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil
arXiv AI
Aug 28

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

The article examines how a difference‑in‑differences (DiD) analysis on a censored rating scale can produce misleading effects. It demonstrates that each DiD component is censored by its own share, causing differential attenuation that can fabricate an interaction effect when the two responses are unequally censored. Using a pre‑registered audit of an LLM judge, the authors show that the reported significant interaction is largely an artifact of this censoring mechanism, with the true preference effect being null.

By Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang