arXiv AI By Donghyun Lee (Dongguk University), Juntae Kim (Dongguk University)

A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG

Read the original on arXiv AI →

arXiv:2607. 03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 16

Stream Assembly Is an Uncontrolled Treatment in Streaming Intrusion-Detection Benchmarks

The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.

By Michel A. Youssef
arXiv AI
6d ago

Depth, Not Breadth: Best-of-N Jailbreaking Beyond Surface Noise

The paper investigates how allocating a query budget to structural depth rather than surface variation improves jailbreak success against the SAGE self‑check defense. By using a best‑of‑N approach over a code‑completion encoding, the authors achieve 67%, 22%, and 15% success rates on three open‑weight targets—far exceeding the 4.7% and 3.0% rates of single‑draw encoding and character‑search methods. The study demonstrates that depth of encoding and breadth of variation independently undermine transform and gate defenses, and that repeated sampling can inflate perceived robustness.

By Haoyu Zhang, Hanwen Liu, Yang Chen, Shibo Zheng, Xiangchen Guan, Zhuoxi Wang, Zijian Xiao, Xiao Luo, Yi Feng, Haowen Xu, Mohammad Zandsalimy, Shanu Sushmita
arXiv AI
Jun 3

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

arXiv:2606. 02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat.

By Alexandre Cristov\~ao Maiorano
arXiv Machine Learning
Jul 30

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

arXiv:2607. 26574v1 Announce Type: cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap.

By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita