arXiv Machine Learning

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.

arXiv AI
Aug 7

When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents

arXiv:2608. 05810v1 Announce Type: new Abstract: Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not monotonic: past a critical pool size, newly added skills degrade performance instead of improving it.

By Linfang Shang, Ming Xu, Yiding Sun, Tianle Xia, Lingxiang Hu, Lan Xu, Ning Zheng
arXiv AI
Sep 17

Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models

The study investigates whether unusually high‑gain parameters in transformer models—specifically gated feed‑forward network (gated‑FFN) rows—play a critical functional role across both text and genomic foundation models. By computing exact bilinear weight operators and testing structural extremeness, the authors find that high‑gain rows are enriched for functional importance but do not reliably predict causal effect size or severity. The analysis reveals model‑specific causal organizations, including super‑additive interactions in DNABERT‑2 and position‑localized dependencies in GENERator, indicating that structural prominence signals enrichment rather than calibrated criticality.

By Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares
arXiv Machine Learning
1d ago

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

The paper investigates how fine‑tuning large language models with a small number of harmful examples can erode their refusal behavior, and explores whether localizing safety‑related behavior to specific layers or directions can provide robust defenses. Experiments across six checkpoints from four model families show that harmful and benign prompts remain linearly separable after attack, and that patching clean hidden states or freezing layers up to a transition depth can restore refusal. However, attackers can bypass these defenses by spreading updates or targeting singular directions, indicating that adaptive fine‑tuning can defeat localized repairs and highlighting the need for multiple defensive checks.

By Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim