arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2607. 06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicious are audited or deferred.
By Lucas Pinto
arXiv:2607. 01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable.
By Tung-Ling Li, Hongliang Liu, Yuhao Wu
arXiv:2608. 01676v1 Announce Type: cross Abstract: Sparse attention is widely deployed in long-context serving stacks, yet no framework audits how discarding blocks changes the influence of specific content on model output.
By Xingyu Ren, Youran Sun, Chugang Yi, Haizhao Yang
The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.
By Sahil Kadadekar
arXiv:2609.15017v1 Announce Type: cross
Abstract: Prompt-injection detectors are typically evaluated using aggregate F1 on in-distribution test data, which offers limited insight into behavior under...
By Yusuf Khalid Shire, Sang-Chul Kim
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.
By Yuchen Han, Cheng Yan, Wuyang Zhang
arXiv:2607. 03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning.
By Donghyun Lee (Dongguk University), Juntae Kim (Dongguk University)
arXiv:2608. 16003v1 Announce Type: new Abstract: Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer.
By Parsa Mazaheri, Kasra Mazaheri
arXiv:2607. 09306v3 Announce Type: replace-cross Abstract: Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation).
By Kwan Soo Shin
arXiv:2606. 02276v1 Announce Type: cross Abstract: Vision-language models (VLMs) trained on paired chest radiographs and radiology reports learn a shared embedding space that can preserve instance-level image-report correspondence.
By Soroosh Tayebi Arasteh, Mahshad Lotfinia, Sven Nebelung, Daniel Truhn