← Back to all news
arXiv Computation and Language September 21, 2026 By Maciej Skorski

Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 9

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

arXiv:2601. 21433v2 Announce Type: replace Abstract: Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change in framing.

By Katherine Elkins, Jon Chun
llmssafety
More like this →
arXiv AI
Jun 11

Every Act Has Its Price: Compressed Moral Composition in Frontier LLMs

arXiv:2606. 11232v1 Announce Type: cross Abstract: Existing LLM moral benchmarks usually ask which isolated moral act, value, or foundation a model prefers.

By Weijia Zhang, Ruiqi Chen, Yunze Xiao, Weihao Xuan
llmsbenchmarks
More like this →
arXiv Machine Learning
Sep 10

Decomposing LLM-Judge Uncertainty to Target Expert Labels

arXiv:2609.06444v2 Announce Type: cross Abstract: An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties:...

By Ryan Lail
llms
More like this →
arXiv AI
Aug 14

Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments

arXiv:2608. 12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs).

By Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik, Mateja Centa Strahovnik, Jonas Miklav\v{c}i\v{c}, Marko Robnik \v{S}ikonja
llmsagentsbenchmarkssafety
More like this →
arXiv AI
Aug 17

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

arXiv:2608. 14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.

By Dipankar Sarkar
llmsbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 30

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

arXiv:2510. 06096v3 Announce Type: replace Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge.

By Matthieu Bou, Nyal Patel, Arjun Jagota, Satyapriya Krishna, Sonali Parbhoo
llmsreinforcement-learningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea