arXiv AI

Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing

arXiv:2607. 28814v1 Announce Type: cross Abstract: In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy).

arXiv Machine Learning
Sep 22

Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch

The paper demonstrates that a preference‑optimization objective can learn to distinguish reliable from unreliable sources by installing a prior‑dependent reliability switch. By training on data where a source’s stated reliability is paired with its answer, the model learns to flip its response only when the stated reliability exceeds a threshold that grows with the model’s prior. Experiments on Qwen2.5‑7B‑Instruct and Llama‑3.1‑8B show that this switch generalizes to unseen reliability values and follows stated reliability over role prestige, whereas supervised imitation fails to learn it.

By Sen Yang, Yuen-Hei Yeung
arXiv AI
2d ago

RISED: RubrIcs for agentic multi-environment Selection and sElf-Distillation

RISED introduces a framework that uses rubric-based textual feedback to improve training of a single large language model (LLM) agent across multiple interactive environments. By having an LLM judge tag rollouts with a shared rubric vocabulary, the system guides both online data selection and policy supervision, enabling richer cross‑environment relationships and within‑group reward contrast. Experiments show that RISED achieves the highest mean pass rate and ranks first or second in every individual environment, with rubric analysis revealing behavioural changes behind these gains.

By Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu
arXiv Machine Learning
Sep 17

No Usable Linear "Capitulation Direction" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback

The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.

By Saad Aamir, Muhammad Awais Bin Adil
arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
arXiv AI
Jun 30

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

arXiv:2606. 30627v1 Announce Type: cross Abstract: Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-supported behaviour, the argument goes, it is less likely to exploit imperfections in a learned reward model.

By Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary
arXiv Computation and Language
Sep 7

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

The paper introduces a boundary-aware self‑distillation framework for controlled large language model safety refusal, addressing the need for different refusal boundaries within the same topic. It combines controlled topic generation, coverage repair, in‑distribution compensation data, and harmful‑benign pairs to train and evaluate refusal behavior. Experiments on Qwen3‑8B show that escalating retries dramatically improve target‑domain refusal rates while reducing unsafe responses, though they also increase over‑refusal, highlighting the trade‑off between safety and usability.

By Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us
arXiv AI
Jun 30

Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents

arXiv:2606. 30383v1 Announce Type: new Abstract: A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results) while also conversing in a separate channel with a counterparty whose interests may diverge (negotiating with a vendor, screening inbound requests, or mediating between employees).

By Bojie Li, Noah Shi
arXiv Machine Learning
Jun 5

Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability

arXiv:2605. 03217v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in settings that require nuanced ethical reasoning, yet existing bias evaluations treat model outputs as simply "biased" or "unbiased.

By Yash Aggarwal, Atmika Gorti, Vinija Jain, Aman Chadha, Krishnaprasad Thirunarayan, Manas Gaur