HoneyRoute is an inference‑serving layer that identifies malicious requests and redirects them to a honeypot model, protecting the production model while collecting adversarial interactions for intelligence. It combines a lightweight streaming router, a dual‑implementation honeypot, and an analysis loop that turns captured interactions into attacker fingerprints for retraining. In evaluations on production traces and a seven‑domain attack corpus, HoneyRoute achieves high detection F1 scores with minimal latency overhead, significantly reduces token consumption during flooding attacks, and improves fidelity‑traceability trade‑offs compared to naive bait injection.
By Han Jin
arXiv:2609.39549v1 Announce Type: cross
Abstract: As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant secu...
By Zezhong Wang, Xueyang Tang, Rui Lian, Yang Lou, Heqing Huang
arXiv:2607. 07368v1 Announce Type: cross Abstract: AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent.
By Oliver Makins, Orazio Angelini, Zohreh Shams, Mary Phuong
arXiv:2606. 06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately.
By Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad IV, Joachim Schaeffer, Ram Potham, Tyler Tracy
arXiv:2606. 13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross.
By Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang
arXiv:2608. 15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive.
By Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li