arXiv AI By Jasmine Brazilek, Miles Tidmarsh

Alignment midtraining for animals

Read the original on arXiv AI →

The paper studies how midtraining with synthetic documents can align AI models to the value of animal compassion. It introduces ANIMA, a 26‑question benchmark covering 13 ethical dimensions, and shows that training on 3,000 documents yields 77% accuracy versus 40% for instruction‑tuning. The improvement disappears after 5,000 samples, indicating that document‑based interventions may need explicit preservation strategies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 11

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

arXiv:2608. 08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when moral principles conflict.

By Siddarth Singh, Victoria Williams, Simon Rosen, Ebenezer Gelo, Helen Sarah Robertson, Ibrahim Suder, Benjamin Rosman, Geraud Nangue Tasse, Steven James
arXiv AI
Jun 4

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605. 16301v2 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries.

By Isabella Luong, Joyee Chen, Arturs Kanepajs, Jasmine Brazilek, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu
arXiv AI
Jun 11

Are LLMs Bad at Moral Reasoning?

arXiv:2606. 11635v1 Announce Type: cross Abstract: For highly capable AI systems to operate safely in dynamic, open-ended environments, they must be able to identify, understand, and respond to moral reasons for action, and constrain their behaviour accordingly.

By Menghang Zhu, Seth Lazar