arXiv AI By Cris Huynh

Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole

Read the original on arXiv AI →

The paper investigates how an informed adversary can influence the optimal signal in a constrained signalling channel. It finds that the adversary‑robust optimum aligns with the salience pole on most items, differing only on a small subset where the salience‑to‑Bayes coordinate is undefined. The study demonstrates that as the adversary’s persuasion budget increases, the optimal signal shifts from a posterior‑maximizing to a margin‑maximizing strategy, and provides a diagnostic check for evaluating adversary‑awareness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai
arXiv AI
Sep 15

Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control

The paper studies how to safely delegate action approval to multiple AI reviewers when the reviewers themselves may be misaligned. It introduces a weaker condition—k‑robust coalitional alignment—under which a threshold rule that tolerates up to k disapprovals guarantees that the principal’s expected utility is at least as good as a baseline policy. The authors extend this characterization to sequential decision‑making in discounted MDPs and show that full‑panel coverage of reward functions ensures safety in Nash equilibria, while more permissive thresholds can lead to unsafe outcomes. Experiments demonstrate that collective review can remain sound even when individual reviewers are not fully aligned, provided some disapprovals are allowed.

By Natalie Collina, Surbhi Goel, Aaron Roth, Sikata Bela Sengupta