The paper evaluates the common assumption that combining flow statistics and TLS handshake fingerprints improves encrypted command-and-control detection. Using 17,577 TLS flows from 62 real Cobalt Strike captures, the authors show that data leakage and preprocessing choices inflate performance metrics, revealing that the true benefit of multi-view fusion is minimal (0.022 F1). They also uncover that many captures contain only benign traffic and that class imbalance is an artifact of analysis rather than a real feature of the task.
By Hoang-Huy Nguyen-Huu, Van-Tri Phan, Khuong Nguyen-An
MoSign is a challenge-response authentication system that embeds a time‑varying keyed message into the motion of virtual‑reality users, allowing them to prove identity while keeping their avatars anonymous. The watermark is added to the latent space of a motion variational autoencoder and is provably indistinguishable from unwatermarked motion, with security tied to breaking a pseudorandom function. Experiments on HumanML3D and BOXRR‑23 show high authentication accuracy, low false‑accept rates, and resilience against realistic recapture attacks while remaining undetectable by standard detectors.
By Xujun Che, Thomas Carr, Depeng Xu, Aidong Lu, Shuhan Yuan
TwinMark is a watermarking scheme that embeds a single SHAKE128 secret into a vision model using two complementary linear functionals of model-output summaries: a covariance projector (cov‑Feat) and a class‑conditional Fisher‑aligned linear carrier (cc‑FALC). These readouts cover both classifier APIs attacked by KL knowledge distillation and representation‑only hosts attacked by feature‑matching distillation, each providing a teacher‑measurable a posteriori certificate. Across 13 attacks on datasets such as CIFAR‑10, CIFAR‑100, and Mini‑ImageNet, TwinMark remains detectable on every post‑attack model that retains task utility, survives cross‑architecture distillation onto ResNet‑18/50, VGG‑16, and MobileNet‑V3, and can be ported to GNSS few‑shot, VOC detection, ISIC segmentation, and STL‑10 SimCLR.
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2606. 10456v1 Announce Type: cross Abstract: AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps, each individually below any per-step alarm.
By Zhang Qinqin, Gao Yuze
arXiv:2606. 29441v1 Announce Type: cross Abstract: Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists.
By Subhadip Mitra
The paper investigates whether passive motion traces recorded during selfie capture can serve as an auxiliary signal for detecting spoofing and verifying users in mobile remote identity verification systems. It introduces the CanSelfie dataset, comprising 375 multi‑sensor sequences from 30 participants, and evaluates seven time‑series classifiers and eight anomaly detectors across various sensor configurations. Results show that accelerometer‑only classifiers achieve very low false rejection rates, while certain models achieve low false acceptance rates and high verification accuracy, indicating that selfie‑capture motion is a promising low‑friction evidence channel.
By Erkka Rantahalvari, Olli Silv\'en, Zinelabidine Boulkenafet, Constantino \'Alvarez Casado
The paper reports a privacy breach in a two-node split‑LLM training system where the returned gradient reveals which data rows were real, despite the system passing standard privacy checks. By exploiting the fact that decoy rows produce zero gradients, an attacker can identify real rows with 100% accuracy across multiple runs. The authors demonstrate that adding gradient clipping and noise can mitigate the leak, but the system remains vulnerable to several untested attack vectors.
By Georgios Politis, Evangelos Pappas
DriftNet is a dual‑head trajectory Transformer designed to detect and localize prompt injection attacks in large language model agents. It processes logged tool‑call trajectories, classifying each as compromised or not while labeling every step as benign, injection point, hijacked, or failed injection. On the AgentDrift benchmark, DriftNet achieves high accuracy, with an F1 score of 0.983, 98.7% exact injection‑point recovery, and low false‑alarm rates.
By Asif Pinjari, Mithun Paul Saint-Germain
The paper presents a preprocessor that recovers and decodes encoded content in vision‑language models to close the decode gap that allows harmful requests to bypass safety classifiers. Evaluated against eleven encoding attacks, the preprocessor raises block rates from 0 % to 67‑90 % but also increases benign over‑refusal, and no configuration achieves an ensemble attack‑success rate below 40 % while keeping benign over‑refusal under 70 %. The study shows that closing one encoding channel merely relocates success rather than eliminating it, highlighting the limits of recovery‑based defenses.
By Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Hanwen Liu, Yi Feng, Haowen Xu, Xiangchen Guan, Yang Chen, Zijian Xiao, Xiao Luo, Mohammad Zandsalimy, Shanu Sushmita
arXiv:2609.17856v1 Announce Type: new
Abstract: Heterogeneous cooperative perception (CP) enables connected vehicles with diverse sensor setups to share spatial awareness via compact feature maps, wh...
By Chenyi Wang, Yutong Liu, Qingzhao Zhang, Ming F. Li
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense paradigms (no defense, static steering, CAST, AlphaSteer, probe-gated) across seven instruction-tuned models (7-31B) and five attack types (GCG, AutoDAN, DeepInception, prefilling, intent laundering).
arXiv:2607. 11751v1 Announce Type: cross Abstract: As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own.
By Yibo Hu, Ren Wang