Now You (Still) See Me: Detecting Evasive Steganographic Payloads in LLMs
arXiv:2606. 09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs.
The paper studies how small, targeted changes to model parameters affect classical machine learning models (HMM and SVM) versus deep learning models (MLP and LSTM). Using the Drebin Android malware dataset, it finds that classical models are brittle, with a few parameter changes dramatically altering behavior, and have limited steganographic capacity. In contrast, neural networks are parameter‑redundant, allowing many parameters to be altered with minimal impact, thus supporting higher steganographic capacity.
arXiv:2606. 09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs.
arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.
arXiv:2607. 20003v1 Announce Type: cross Abstract: An increase in advanced Android malware requires the use of deep learning models, which can run on Android devices.
arXiv:2507. 01752v4 Announce Type: replace-cross Abstract: Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation.
arXiv:2607. 19894v1 Announce Type: cross Abstract: Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs.
arXiv:2505. 03646v5 Announce Type: replace-cross Abstract: Adversarial robustness of deep autoencoders (AEs) has received less attention than that of discriminative models, although their compressed latent representations induce ill-conditioned mappings that can amplify small input perturbations and destabilize reconstructions.
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.
arXiv:2505. 19840v3 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) have achieved widespread success yet remain prone to adversarial attacks.
arXiv:2608. 15113v1 Announce Type: cross Abstract: Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings.
The paper introduces Replicant, a deep reinforcement learning framework that learns to evade malware detectors under a strict label‑only black‑box threat model. Replicant generates reusable policies for modifying malware samples and deciding when to query the target, and it transfers across different samples, detectors, and feature spaces. In experiments on seven Android malware detectors and three feature spaces, Replicant achieves a mean attack success rate of 78.8%, outperforming state‑of‑the‑art methods by 20.9%–39.2% and providing a stronger signal for adversarial training to harden detectors.
arXiv:2406. 10090v3 Announce Type: replace Abstract: Thanks to their extensive capacity, over-parameterized neural networks exhibit superior predictive capabilities and generalization.
arXiv:2609.36117v1 Announce Type: new Abstract: Securing modern AI systems against backdoor attacks remains an open challenge and requires fundamentally principled estimates of the adversary's budget...