arXiv:2608.21570v1 Announce Type: new
Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guard models hol...
By Edson Rodrigues da Cruz Filho, Paulo Ricardo Ferreira Neves, Paulo Henrique Eleuterio Falsetti, Jo\~ao Vitor Pavan, Ian Degaspari, Henrique Vieira Laturrague, Patrick Vieira Laturrague, Guilherme Nielsen Dias, Marccello Wilson Perez Berto, Gustavo Voltani Von Atzingen
The paper investigates whether stacking multiple defenses around large language models (LLMs) truly compounds security. Using the Adversary Access‑Tier Model (AATM) and a cost‑tiering system, the authors analyze a seven‑layer defense stack and find that failure correlations between layers are consistently positive, meaning the residual attack success is higher than the multiplicative prediction. Despite high coverage and low false refusals, the stack’s performance is largely driven by common architectural causes rather than diverse, independent defenses.
By Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed
The paper investigates learning with monotone adversarial corruptions, extending previous binary classification results to multiclass and partial binary settings. It shows that even a small number of strategically inserted corrupted examples can render a learnable multiclass problem with DS dimension 2 completely unlearnable, and provides matching upper bounds when the adversary’s budget is sublinear. The work also demonstrates that classic error rates remain attainable under bounded or limited‑view adversaries.
By Julian Asilis, Shaddin Dughmi, Chirag Pabbaraju
arXiv:2605.25663v2 Announce Type: replace-cross
Abstract: Black-box adversarial attacks that minimize only the ground-truth confidence suffer from class drift: perturbations wander through the featur...
By Florent Tariolle, Florian Yger
arXiv:2609.39338v1 Announce Type: new
Abstract: Knowledge distillation transfers knowledge by encouraging a student to match a teacher's predicted class probabilities. These probabilities express not...
By Qianfeng Yuan, Wenbing Tao
The paper investigates the impact of label‑flipping attacks on distributed machine learning, where an adversary can only flip a limited number of training labels. It formalizes the attack as a per‑round constrained optimization problem, derives a greedy label‑selection rule for logistic regression, and shows that this rule is provably optimal under mean aggregation. Experiments demonstrate that optimized label flipping can significantly degrade model accuracy, outperforming random flips, and that the attack transfers to other robust aggregators such as coordinate‑wise median and trimmed mean.
By Abdessamad El-Kabid, El-Mahdi El-Mhamdi