A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG
arXiv:2607. 03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning.
The paper demonstrates that the way evaluation streams are assembled in streaming intrusion‑detection benchmarks—by interleaving, pooling, or replaying network captures—acts as an uncontrolled experimental variable that can significantly alter performance metrics. In the CICIDS2017 benchmark, reordering the same set of records under a fixed split changes the held‑out samples’ overlap, prevalence, and even reverses the ranking of two deterministic scorers. Similar effects are observed in the LITNET‑2020 benchmark, where pooling disjoint captures yields a single operating point that masks large variations in per‑capture prevalences, and minor changes in batch composition can shift reported AUC‑PR values by a few thousandths.
arXiv:2607. 03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning.
arXiv:2609.25052v1 Announce Type: new Abstract: "An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to...
arXiv:2605. 24696v2 Announce Type: replace-cross Abstract: Streaming intrusion detection systems must process flows continuously under bounded memory, yet most leave alerting-threshold selection as a post-hoc tuning problem incompatible with production, where operators commit in advance to alert budgets, misclassification costs, and Service Level Objectives.
arXiv:2609.22818v1 Announce Type: cross Abstract: Memory-poisoning defenses for LLM agents are typically evaluated by their ability to prevent attacks. However, the traffic they process is rarely adv...
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
arXiv:2607. 13203v1 Announce Type: cross Abstract: False alarms remain a major barrier to deploying network intrusion detection systems (NIDS).
arXiv:2608. 15761v1 Announce Type: cross Abstract: Edge-IIoTset is the reference benchmark for machine-learning intrusion detection in the industrial Internet of Things, and results reported on it cluster above 99%.
arXiv:2607. 09800v2 Announce Type: replace Abstract: Master weights and stochastic rounding bypass invisible stored-weight updates but do not locate lost direct-storage proposals or parameters worth protecting.
arXiv:2609.38021v1 Announce Type: cross Abstract: We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, cove...
arXiv:2606. 15474v1 Announce Type: new Abstract: Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a team is paged when the score drifts down.
The paper investigates how to properly validate candidate models before promoting them to replace incumbent classifiers in adaptive network intrusion detection systems. It demonstrates that promotion decisions can be biased by how challengers are constructed and the amount of evidence they receive, and that using self‑contained challenger pipelines and sufficient candidate evidence reduces apparent promotion harm. The study also shows that policy rankings shift with candidate comparability and that no single update policy dominates across benchmarks.
arXiv:2609. 22880v1 Announce Type: cross Abstract: LLM rerankers add of the order of \$0.