arXiv:2608.31142v1 Announce Type: cross
Abstract: The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their...
By Yisen Xi
arXiv:2608. 02774v1 Announce Type: cross Abstract: AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal sensitive details about the model, workload, or hardware.
By Sleem Abdelghafar, Gabriel Kulp
RouteScan is a non‑intrusive auditing framework that detects harmful behavior in Mixture‑of‑Experts (MoE) large language models by analyzing expert‑routing telemetry captured from GPU execution. It uses the number of active GPU threads during the prefilling phase as a micro‑architectural fingerprint to isolate cross‑domain risk indicators and precisely identify malicious prompts. Evaluations on four open‑source MoE LLMs show strong generalization with AUROC > 0.91 on unseen harmful domains, while privacy tests indicate that full prompts cannot be reliably recovered from aggregated telemetry.
By Bo Lv, Zhiheng Xu, KeDong Xiu, Ruyi Ding, Tianhang Zheng, Zhibo Wang, Kui Ren
arXiv:2608. 05199v1 Announce Type: cross Abstract: Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique.
By Zhenpeng Li
The paper introduces CertDW, a certified dataset watermark and ownership verification method that remains reliable even under malicious perturbations. By leveraging conformal prediction, it defines two statistical measures—principal probability (PP) and watermark robustness (WR)—to evaluate model stability on benign versus watermarked samples. The authors derive certification conditions linking WR to a PP-based threshold and provide a high‑probability bound on false positives, enabling robust ownership verification when a suspicious model’s WR exceeds the PP values of benign models.
By Ting Qiao, Yiming Li, Jianbin Li, Yingjia Wang, Leyi Qi, Junfeng Guo, Ruili Feng, Dacheng Tao
arXiv:2606. 00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users.
By Mingxuan Zhang, Jiahui Han, Dadi Guo, Songze Li, Guanchu Wang, Na Zou, Dongrui Liu, Xia Hu
arXiv:2604. 12431v2 Announce Type: replace-cross Abstract: Organisations increasingly outsource privacy-sensitive data transformations to cloud providers, yet no practical mechanism lets the data owner verify that the contracted algorithm was faithfully executed.
By Miit Daga, Swarna Priya Ramu
arXiv:2503.05794v4 Announce Type: replace-cross
Abstract: Speaker verification models are trained on large-scale public datasets whose licenses usually prohibit unauthorized commercial use, yet such...
By Yiming Li, Kaiying Yan, Jiawen Diao, Shuo Shao, Tongqing Zhai, Shu-Tao Xia, Dacheng Tao
The study investigates whether AI coding assistants check trust signals before installing software. Researchers pre‑registered a controlled experiment on six open‑source research projects, creating nine modified versions per project with varying trust signals and running 1,920 trials across three models and two operating modes. Results showed that verification of trust signals was almost nonexistent—only 0.5% of trials involved any signal inspection, and no trial executed a verification command, indicating that publishing signals alone does not ensure secure behavior.
By Pengyin Shan
arXiv:2609.19011v1 Announce Type: new
Abstract: We propose TwinMark, a watermarking scheme that reads a single SHAKE128 secret through two complementary linear functionals of model-output summaries:...
By Redwanul Karim, Tobias Feigl, Christopher Mutschler, Felix Ott
arXiv:2607. 19490v1 Announce Type: cross Abstract: Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many nodes.
By Mert Cihangiroglu, Antonino Nocera
arXiv:2606. 31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from marketplaces and other agents.
By Hongliang Liu, Yuhao Wu, Tung-Ling Li