arXiv:2606. 12703v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) agents increasingly run with persistent memory that accumulates across user sessions.
By Tarun Sharma
The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.
By Michel A. Youssef
The paper investigates how allocating a query budget to structural depth rather than surface variation improves jailbreak success against the SAGE self‑check defense. By using a best‑of‑N approach over a code‑completion encoding, the authors achieve 67%, 22%, and 15% success rates on three open‑weight targets—far exceeding the 4.7% and 3.0% rates of single‑draw encoding and character‑search methods. The study demonstrates that depth of encoding and breadth of variation independently undermine transform and gate defenses, and that repeated sampling can inflate perceived robustness.
By Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Xiao Luo, Yi Feng, Haowen Xu, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2606. 02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat.
By Alexandre Cristov\~ao Maiorano
arXiv:2609.00470v1 Announce Type: new
Abstract: Retrieval-Augmented Generation (RAG) grounds large language models in external corpora, but implicit trust in retrieved documents creates a critical at...
By Muhaimin Bin Munir, Akib Jawad Ononto, Nazia Shehnaz Joynab, Bhavani Thuraisingham, Latifur Khan
arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita