arXiv AI By Alizishaan Khatri, Dun Li Chan

Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families

Read the original on arXiv AI →

arXiv:2608. 08029v1 Announce Type: cross Abstract: Khatri et al.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.