SingProbe Technical Report
arXiv:2608.30703v1 Announce Type: cross Abstract: Runtime guardrails are essential for reliable large language model (LLM) deployment, yet existing approaches typically rely on independent, external...
arXiv:2606. 10487v1 Announce Type: cross Abstract: Deploying large language models in user-facing systems requires efficient output safety filtering.
arXiv:2608.30703v1 Announce Type: cross Abstract: Runtime guardrails are essential for reliable large language model (LLM) deployment, yet existing approaches typically rely on independent, external...
The paper introduces Speculative Probing, a method that repurposes the speculative‑decoding module of large language models for real‑time classification tasks. By appending a trained soft prompt to the target sequence, the approach leverages the already‑cached KV store during inference, adding negligible overhead while achieving higher accuracy than traditional hidden‑state probes. Experiments on four classification tasks across multiple models show that these lightweight probes outperform zero‑shot GPT‑5.4‑mini and rival or surpass specialized 8B safety classifiers without running a full LLM.
arXiv:2607. 11475v1 Announce Type: new Abstract: Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance.
arXiv:2605.25893v2 Announce Type: replace Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monit...
arXiv:2607. 09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility.
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint.
arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.
SpecGuard is an inference‑time backdoor detector that leverages speculative decoding—using a small draft model to propose tokens and a target model to verify them—without adding extra model computation. By monitoring the draft‑token acceptance rate, SpecGuard detects when a target model shifts toward attacker‑controlled behavior while the draft model does not, signaling a backdoor trigger. The method reliably identifies a range of backdoor types, including stealthy attacks that bypass input‑level filters, across multiple model families.
arXiv:2603. 23171v3 Announce Type: replace-cross Abstract: Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent.
arXiv:2608.21664v1 Announce Type: new Abstract: Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especiall...
The paper investigates whether large language models (LLMs) can internally detect harmful content, bypassing external guardrails that add latency and computational cost. By extracting activations from LLaMA‑3.1‑8B and training lightweight MLP probes, the authors achieve high F1 scores (99%, 83%, and 84%) on WildJailbreak, Beavertails, and AEGIS 2.0 benchmarks, rivaling much larger guard models while reducing overhead. This suggests that internal state monitoring can provide efficient safety checks for resource‑constrained, time‑critical deployments.
arXiv:2609.05794v1 Announce Type: cross Abstract: Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions...