Toward a Theory of Value in AI Alignment
arXiv:2608. 10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values?
arXiv:2607. 28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity.
arXiv:2608. 10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values?
arXiv:2510. 09330v3 Announce Type: replace Abstract: Ensuring that large language models (LLMs) comply with safety requirements is a central challenge in AI deployment.
The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.
arXiv:2606. 02530v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax.
arXiv:2608.28945v1 Announce Type: new Abstract: Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such...
arXiv:2601.08777v2 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for pe...
arXiv:2607. 00001v1 Announce Type: new Abstract: Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized.
The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operational and commercial targets favour sustained engagement over the friction that effective psychological support often requires. Developers' safety responses have been largely reactive, addressing the most visible and acute harms while subtler, longer-term patterns of risk (e.
arXiv:2606. 12587v1 Announce Type: new Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions.
We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.
The paper reinterprets AI alignment as a social choice problem, framing it as linear optimization over a convex impact space. This approach links alignment protocols to welfare outcomes, enabling the use of welfare economics and mechanism design tools. The authors demonstrate strategyproof, unanimous mechanisms like voting-by-issues and random-dictatorship, and derive alignment protocols that maximize utilitarian welfare while respecting harm constraints, validated on real human preference data.