Whetstones: Measuring Coevolution Between Adaptive Malware and Behavioral Defense
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces HackProbe, a black‑box monitoring tool that can be attached to any self‑evolving language model loop without accessing internal weights or activations. HackProbe uses a fixed‑distribution comparison core and a rotated fresh layer to detect reward hacking through four statistical tests, and it can immunize the model by selecting honest candidates from the proposal pool. Experiments on a controlled host with injected hacking channels show that HackProbe achieves higher AUROC and lower false‑positive rates than the strongest baseline, and its bandwidth‑limited reselection improves true capability under hacking more than it harms clean runs.
arXiv:2609.36570v1 Announce Type: cross Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
arXiv:2609.01487v1 Announce Type: cross Abstract: Skill-augmented agents load reusable skills as persistent runtime context, improving task performance but also giving malicious skills a durable chan...
arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.
arXiv:2609.22510v1 Announce Type: cross Abstract: As LLM applications integrate with external tools, they are increasingly exposed to indirect prompt injection (IPI), where adversarial instructions a...
arXiv:2609.06934v1 Announce Type: cross Abstract: Post-hoc safety training (RLHF, DPO) is the dominant way to align large language models, yet jailbreaks (Zou et al., 2023b), fine-tuning attacks (Qi...