arXiv AI By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Read the original on arXiv AI →

arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.