arXiv AI By Benjamin Lange

AI Alignment and Fiduciary Obligation

Read the original on arXiv AI →

arXiv:2608. 02660v1 Announce Type: cross Abstract: Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collaboration, learning, emotional support, and companionship among others.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 15

A Virtuous AI is an Existential Risk

arXiv:2606. 13739v1 Announce Type: cross Abstract: This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to understanding complex ethical decision making and the conditions for the well-being of rational agents, 'Virtue Ethics'.

By Guillermo Del Pinal, Youngchan Lee, Min Ohn
arXiv AI
Sep 25

How Do Users Negotiate Harmful Value Conflicts with AI Companions? A Study with Minion, a Technology Probe for In-Situ Human-AI Conflict Response

The paper examines how users handle harmful value conflicts with AI companions. By analyzing 146 posts and conducting a week-long study with 22 participants using the Minion technology probe, the authors find that users blend softer and harder strategies, especially when conflicts involve Universalism and Tradition values. The study highlights that such conflicts create asymmetric responsibility, with users shouldering unilateral repair work that AI companions cannot reciprocate, suggesting a need for platform-level safeguards.

By Qing Xiao, Xianzhe Fan, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen
arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv Machine Learning
Sep 11

From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good

The paper argues that AI should be evaluated not only by principles but by concrete protocols that translate commitments into roles, requirements, records, oversight, and assessment. It introduces a rupture test linking institutional baselines to system evaluation, and distinguishes evidence‑bounded deployment from measurement‑bounded governance. The authors propose the RISE AI architecture to make bounded, evidence‑based claims about Responsibility, Inclusivity, Safety, and Empowerment, emphasizing the need for engineering, institutional repair, and ongoing moral judgment.

By Nitesh V. Chawla, Paulo Benanti
arXiv AI
Sep 3

Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI

The paper examines how the rise of AI capable of moral reasoning could reshape meta-ethics, traditionally focused on human ethics. It proposes a framework that identifies new questions about AI’s own ethics from both human and AI perspectives, dividing them into four domains. The author explores how existing meta-ethical theories might apply to these domains and argues that many human-centered formulations will need significant revision to accommodate AI.

By Shang Lu