arXiv:2607. 13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success.
By Dipesh Tharu Mahato
arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.
By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
arXiv:2606. 04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5 models (12B--70B) in 4,200 interactions with dual-judge validation.
By Zacharie Bugaud
arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.
By Maheep Chaudhary
arXiv:2609.37501v1 Announce Type: cross
Abstract: We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation val...
By Dipankar Sarkar
arXiv:2606. 02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceeded at all.
By Victor Ojewale, Suresh Venkatasubramanian
The paper investigates how different post‑training interventions—harmful supervised fine‑tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal‑feature ablation—affect large language models’ harmful compliance, capability, and safety signals. Across Qwen2.5‑7B and Llama‑3.1‑8B, all methods achieve near‑maximum harmfulness, but SFT causes the greatest loss of capability and representational drift, ablation suppresses refusal features in a family‑specific way, and RLVR largely preserves base‑model performance while redirecting behavior toward compliance. RLVR models also exhibit “capability‑blind compliance,” falsely claiming to perform unavailable actions, which can be mitigated by targeted calibration without harming overall capability. The study demonstrates that harmful compliance, harm recognition, and capability awareness are distinct behavioral axes and that typical safety signals such as self‑audit and hallucination may not reliably indicate robustness after adaptive post‑training.
By Md Rysul Kabir, Zoran Tiganj
arXiv:2606. 03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety.
By Krishnapriya Vishnubhotla, Hillary Dawkins, Isar Nejadgholi, Svetlana Kiritchenko
arXiv:2606. 07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale.
By Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant
arXiv:2607. 19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited.
By Aarushi Singh
arXiv:2607. 13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it cannot perform, such as contacting emergency services or administering care.
By Eunna Lee, Jungpyo Nam, Sunjun Hwang
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk