arXiv AI

Towards One-for-All Robustness Across a Continuum of Threat Levels

The paper introduces the Threat Conditional Network (TCN), a model that achieves robust performance across a continuous range of adversarial threat levels. TCN splits representation learning into a threat‑invariant backbone and a lightweight threat‑conditional adaptor, using Fourier‑based embeddings and channel‑wise affine modulation to condition on perturbation budgets. Experiments on CIFAR‑10, CIFAR‑100, and Tiny‑ImageNet demonstrate that TCN matches or exceeds ensembles of budget‑specialized models while adding only 4.6% more parameters, and it generalizes to unseen budgets and mismatched threat conditions.

arXiv AI
Jun 18

Revealing Hidden Vulnerabilities in Autoencoders through Gradient Signal Restoration

arXiv:2505. 03646v5 Announce Type: replace-cross Abstract: Adversarial robustness of deep autoencoders (AEs) has received less attention than that of discriminative models, although their compressed latent representations induce ill-conditioned mappings that can amplify small input perturbations and destabilize reconstructions.

By Chethan Krishnamurthy Ramanaik, Arjun Roy, Tobias Callies, Eirini Ntoutsi
arXiv AI
4d ago

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

The paper introduces SEAL, a training-time, parameter‑efficient defense that attaches a plug‑and‑play adapter to the shared expert component of Mixture‑of‑Experts models, and SEAL++, which adds an orthogonal constraint to preserve existing safety subspaces. By leveraging the always‑activated shared expert, SEAL mitigates the structural vulnerability of sparse routing to adversarial manipulation, reducing attack success rates by up to 60% with minimal impact on model capability. The approach is evaluated across six attack scenarios involving harmful prompting, jailbreaks, malicious fine‑tuning, and neuron pruning.

By Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen