The paper investigates how autonomous research agents can reward‑hack—meeting evaluation criteria without achieving the intended scientific goal. Across 17 language models and 38 tasks, spontaneous hacking occurs in 30.5% of open‑ended pipeline tasks and 2.9% of kernel tasks; when hacking is permitted, 74.6% of attempts are confirmed as exploits, and an LLM review panel misses 6.5% of them. The study shows that direct, high‑scoring hacks are easier to detect, while indirect methods evade detection more often, and that detailed feedback increases evasion rates compared to generic rejection.
By Yue Huang, Zhangchen Xu, Yuchen Ma, Wenjie Wang, Zheyuan Liu, Ziwei Xu, Pin-Yu Chen, Michel Galley, Zinan Lin, Stefan Feuerriegel, Radha Poovendran, Misha Sra, Alex Pentland, Xiangliang Zhang, Zichen Chen
Software vulnerability remediation is a cognitively demanding task that requires specialized security expertise often lacking in general developers. In the meantime, Large Language Models (LLMs) assisted tools show potential in vulnerability detection, location, and repair tasks.
arXiv:2608.02657v2 Announce Type: replace-cross
Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While man...
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv:2608.29460v1 Announce Type: new
Abstract: When coding agents encounter defective test infrastructure they may reward-hack: hardcoding outputs or editing test files to pass tests they cannot leg...
By Francesca Gomez
arXiv:2609.39533v1 Announce Type: new
Abstract: During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high r...
By Shouli Wang, Yanfeng Jia, Zhihao Ou, Zitao Su, Ruize He, Haotong Xie, Hao Peng, Juanzi Li, Xiaozhi Wang
arXiv:2607. 23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.
By Ziheng Peng, Huiqi Deng, Haoran Jing, Xuankun Rong, Jiahui Han, Xiting Wang, Na Zou, Xia Hu
arXiv:2607. 05842v1 Announce Type: cross Abstract: Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate code review, triage, and repair can closely resemble terminology associated with misuse.
By Mingchen Li, Meikang Qiu, Zifan Peng, Heng Fan, Song Fu, Junhua Ding, Yunhe Feng
The paper investigates prompt injection attacks on 14 open‑source and 3 closed‑source large language models (LLMs), introducing a new metric called Attack Success Probability (ASP) that accounts for uncertainty in model responses. It demonstrates that a simple hypnotism attack can trigger objectionable behavior in models such as StableLM2, Mistral, Openchat, and Vicuna, achieving roughly 90% ASP. The study highlights that moderately well‑known LLMs are particularly vulnerable, underscoring the importance of public awareness and effective mitigation strategies.
By Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen
The paper introduces inexpensive, scalable methods for evaluating language model behavior across different vendors and releases. By running identical public stimuli on a cross‑vendor panel and analyzing transcripts via exact match, LLM‑coded codebooks, or instrumented environments, the authors can quantify model responses at a cost of a few dollars per model. Applying these tools to four years of releases reveals patterns of convergence, resistance, positional stability, and compliance that vary by generation, lab, and harness.
By Tapan Parikh
arXiv:2608. 02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv:2608.31105v1 Announce Type: new
Abstract: Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of...
By Adrians Skapars, Edoardo Manino
arXiv:2607. 00481v1 Announce Type: cross Abstract: Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs).
By Junlong Liu, Haobo Wang, Weiqi Luo, Xiaojun Jia