arXiv Machine Learning

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

arXiv:2606. 01456v1 Announce Type: new Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sales assistants for purchases, negotiation agents for concessions.

arXiv AI
6d ago

LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information

LAVOIR is a single‑pass decision encoder that not only predicts answers to typed questions but also identifies which missing pieces of information (slots) would most improve its confidence. By placing candidate slots next to answer options, one forward pass yields both the decision distribution and the expected value of asking each slot, without requiring human labels. In controlled experiments, LAVOIR’s question policy matches a greedy oracle and improves accuracy by up to 14.1 points over never asking, while on real conversations it raises accuracy by 8.3 points with minimal questioning.

By Furkan Yilmaz, Habibe Aleyna Tasdemir, Muhammed Faruk Gozay
arXiv AI
Aug 17

Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation

arXiv:2608. 13564v1 Announce Type: new Abstract: Evaluating language-model agents at scale increasingly relies on a second language model as an automatic judge, because the gold signal, an executable environment reward, is expensive, slow, or unavailable at deployment time.

By Darragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
arXiv AI
Jun 2

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

arXiv:2606. 00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts.

By Rana Muhammad Usman
arXiv AI
2d ago

Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole

The paper investigates how an informed adversary can influence the optimal signal in a constrained signalling channel. It finds that the adversary‑robust optimum aligns with the salience pole on most items, differing only on a small subset where the salience‑to‑Bayes coordinate is undefined. The study demonstrates that as the adversary’s persuasion budget increases, the optimal signal shifts from a posterior‑maximizing to a margin‑maximizing strategy, and provides a diagnostic check for evaluating adversary‑awareness.

By Cris Huynh
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
3d ago

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets

The paper introduces PACT, a method for unlearning deceptive behaviors in large language models by using contrastive forget sets that compare a model’s responses under deceptive and neutral contexts. PACT trains the model to produce pressure‑aware counterfactual targets, preserving benign system‑prompt adherence and reasoning traces while dramatically reducing deception rates from over 50% to under 3% on 32B reasoning models.

By Haoran Tang, Rajiv Khanna