Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 15420v1 Announce Type: cross Abstract: A constitution tells a language model what to value, but little tells us whether it does.
arXiv:2609.14759v1 Announce Type: cross Abstract: Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model...
The paper introduces PACT, a method for unlearning deceptive behaviors in large language models by using contrastive forget sets that compare a model’s responses under deceptive and neutral contexts. PACT trains the model to produce pressure‑aware counterfactual targets, preserving benign system‑prompt adherence and reasoning traces while dramatically reducing deception rates from over 50% to under 3% on 32B reasoning models.
arXiv:2608.28648v1 Announce Type: new Abstract: We study how instruction-tuned LLMs arbitrate direct conflicts between system and user instructions. We introduce a benchmark of 41 paired constraints...
arXiv:2609.39702v1 Announce Type: new Abstract: Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choic...
Augur is a synthetic decision laboratory that simulates how users will react to product and policy changes before they are released. It constructs a typed knowledge graph from change documents, populates a persona market, runs simulations, and produces an auditable decision memo recommending one of five actions. Using a dataset of 50 real episodes (Gold‑50), the authors evaluate the system’s five‑way release verdicts and find that evaluation design, rather than model capability, largely drives performance differences among frontier and open‑weight models.