arXiv AI By Ke Miao, Jiaxin Li, Hongliang Chen, Yuke Hu, Zhan Qin

Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

Read the original on arXiv AI →

arXiv:2606. 16808v1 Announce Type: new Abstract: While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.