arXiv AI By Aleks Knoks, Marija Slavkovik

Metanormative Theory for RL-Based Moral Agents

Read the original on arXiv AI →

arXiv:2608. 08220v1 Announce Type: new Abstract: The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with human values and that act in ethically acceptable ways.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment

The paper argues that AI alignment depends on a system’s ability to exhibit a coherent moral policy—stable, monotonic, decisive, and Pareto‑viable—rather than on any specific moral standard. The authors test nine large language models across varied moral scenarios and find that none maintain consistent verdicts, with surface‑form changes causing up to 99% shifts in outcomes. This indicates that current LLM agents lack the structural moral competence required for meaningful alignment.

By Arno Libert, Derck W. E. Prinzhorn, Daan R. Henselmans
arXiv AI
Sep 12

Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents

The paper proposes a developmental framework for autonomous artificial agents that emphasizes learning social norms and alignment through direct interaction with dynamic environments. It argues that intrinsic motivations such as curiosity and competence can guide exploration, but also complicate alignment with human goals. By drawing parallels to child development, the authors suggest that regulatory sandboxes serve as pedagogical spaces where agents gradually acquire moral agency and adapt their behaviors through experience and cooperation.

By Marica Notte, Ludovica Marinucci, Vieri Giuliano Santucci
arXiv AI
Sep 3

Meta-ethics and AI: exploring the novel meta-ethical questions in the era of AI

The paper examines how the rise of AI capable of moral reasoning could reshape meta-ethics, traditionally focused on human ethics. It proposes a framework that identifies new questions about AI’s own ethics from both human and AI perspectives, dividing them into four domains. The author explores how existing meta-ethical theories might apply to these domains and argues that many human-centered formulations will need significant revision to accommodate AI.

By Shang Lu