arXiv AI By Orion Reblitz-Richardson

Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Oct 1

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets

The paper introduces PACT, a method for unlearning deceptive behaviors in large language models by using contrastive forget sets that compare a model’s responses under deceptive and neutral contexts. PACT trains the model to produce pressure‑aware counterfactual targets, preserving benign system‑prompt adherence and reasoning traces while dramatically reducing deception rates from over 50% to under 3% on 32B reasoning models.

By Haoran Tang, Rajiv Khanna
arXiv AI
Sep 25

Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes

Augur is a synthetic decision laboratory that simulates how users will react to product and policy changes before they are released. It constructs a typed knowledge graph from change documents, populates a persona market, runs simulations, and produces an auditable decision memo recommending one of five actions. Using a dataset of 50 real episodes (Gold‑50), the authors evaluate the system’s five‑way release verdicts and find that evaluation design, rather than model capability, largely drives performance differences among frontier and open‑weight models.

By Rahul Khedar, Mayank Malhotra, Avinash Karn