arXiv AI By Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere, Linh Le, David Williams-King, Adam Oberman

Efficient Safety Alignment of Language Models via Latent Personality Traits

Read the original on arXiv AI →

arXiv:2607. 07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 9

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

arXiv:2606. 08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these evaluations target outputs rather than representation-level vulnerability under intervention.

By Enyi Jiang, Anders Gj{\o}lbye, Yibo Jacky Zhang, Sanmi Koyejo