arXiv AI By Najmul Hasan

Honeypot Protocol

Read the original on arXiv AI →

arXiv:2604. 13301v1 Announce Type: cross Abstract: Trusted monitoring, the standard defense in AI control, is vulnerable to adaptive attacks, collusion, and strategic attack selection.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 10

HoneyRoute: Honeypot-Model Routing for Adversarial LLM Serving

HoneyRoute is an inference‑serving layer that identifies malicious requests and redirects them to a honeypot model, protecting the production model while collecting adversarial interactions for intelligence. It combines a lightweight streaming router, a dual‑implementation honeypot, and an analysis loop that turns captured interactions into attacker fingerprints for retraining. In evaluations on production traces and a seven‑domain attack corpus, HoneyRoute achieves high detection F1 scores with minimal latency overhead, significantly reduces token consumption during flooding attacks, and improves fidelity‑traceability trade‑offs compared to naive bait injection.

By Han Jin
arXiv AI
Jun 12

The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems

arXiv:2606. 13079v1 Announce Type: cross Abstract: Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical red lines that frontier AI systems must not cross.

By Jiaqi Luo, Jiarun Dai, Zhile Chen, Jia Xu, Weibing Wang, Yawen Duan, Brian Tse, Geng Hong, Xudong Pan, Yuan Zhang, Min Yang
arXiv AI
Aug 18

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

arXiv:2608. 15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive.

By Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li