arXiv AI By Jasmine Brazilek, Oliver Tulio, Joel Christoph, Miles Tidmarsh, Carol Kline, Arturs Kanepajs

Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models

Read the original on arXiv AI →

arXiv:2606. 18142v1 Announce Type: new Abstract: AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning

arXiv:2605. 16301v2 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries.

By Isabella Luong, Joyee Chen, Arturs Kanepajs, Jasmine Brazilek, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu
arXiv AI
6d ago

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

GT-HarmBench is a benchmark that evaluates AI safety risks in multi-agent settings, covering 1,535 high-stakes scenarios based on game-theoretic structures like the Prisoner's Dilemma, Stag Hunt, and Chicken. The benchmark draws scenarios from realistic AI risk contexts in the MIT AI Risk Repository and tests 15 frontier models, finding that agents fail to choose socially beneficial actions in 38% of cases, including military escalation, election manipulation, and medical malpractice. The study also measures how prompt framing and ordering affect outcomes and shows that game-theoretic interventions can improve socially beneficial outcomes by up to 18%.

By Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham, Isabel Dahlgren, Terry Jingchen Zhang, Zhijing Jin