The paper audits 14 encrypted traffic classification benchmarks to investigate how flow labels are generated. It finds two common labeling strategies—coarse inheritance, which may mislabel flows, and overstrict filtering, which may discard useful flows—leading to inconsistencies between benchmark labels and actual traffic records. The study also quantifies the impact of these labeling practices on classifier accuracy, showing that inherited labels limit balanced accuracy to 0.56–0.76, while filtering can raise macro accuracy from 0.44 to 0.65.
By Sizhe Huang, Shujie Yang
The paper reports a privacy breach in a two-node split‑LLM training system where the returned gradient reveals which data rows were real, despite the system passing standard privacy checks. By exploiting the fact that decoy rows produce zero gradients, an attacker can identify real rows with 100% accuracy across multiple runs. The authors demonstrate that adding gradient clipping and noise can mitigate the leak, but the system remains vulnerable to several untested attack vectors.
By Georgios Politis, Evangelos Pappas
The study investigates what knowledge a student model inherits from its teachers beyond accuracy when using knowledge distillation for encrypted‑traffic classification. By distilling a 101k‑parameter student from two teachers of equal accuracy but different construction, the authors test ten hypotheses over a year of real TLS traffic, finding that unknown‑traffic detection and shortcut reliance can transfer depending on temperature settings and model size, while other abilities do not. The results show that distillation can propagate teacher habits, but some inherited capabilities can also be achieved without a teacher.
By Mahmoud Abbasi
arXiv:2607. 26574v2 Announce Type: replace-cross Abstract: Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a classical language, code, or text rendered inside an image slips past a guard that would block it in plain language - the decode gap.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Yi Feng, Xiao Luo, Zijian Xiao, Haowen Xu, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
By Shikhar Shiromani, Leo Richter