arXiv AI By Aryan Keluskar, Amrita Bhattacharjee, Huan Liu

ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs

Read the original on arXiv AI →

arXiv:2607. 14285v1 Announce Type: cross Abstract: Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict?

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 29

Do Models Fake Alignment Without Clear Consequences?

arXiv:2607. 24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking.

By Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao