arXiv Machine Learning By Yang Gao (Veyon Solutions)

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

Read the original on arXiv Machine Learning →

arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.