arXiv Machine Learning

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

arXiv:2607. 09842v1 Announce Type: new Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state trajectories of four open-weight transformer language models spanning four post-training regimes: no training (Gemma-4-E4B base), multimodal RLHF (Gemma-4-E4B-it), RL distillation (DeepSeek-R1-Distill-Qwen-7B), and SFT (Qwen2.

arXiv Machine Learning
Aug 27

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

The paper investigates why the orthogonal optimiser Muon outperforms Adam in large language model pretraining by analysing the spectral properties of Transformer loss landscapes. It finds that Muon’s momentum buffers exhibit an anisotropic spectral profile with a volatile head and a tolerant bulk, enabling larger effective step sizes. Building on this insight, the authors propose Spectral‑Aware Muon (SAMuon) and a lightweight variant, which adjust the bulk scaling while keeping the head unchanged, achieving 13–24 % fewer training tokens than Muon without extra FLOPs.

By Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland
arXiv Machine Learning
Aug 19

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.

By Sihan Wang, Jiayi Zhao, Qingyan Cao, Hongbo Yao, Lin Shu
arXiv Machine Learning
Sep 11

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients

The paper proposes a flexible four‑coefficient parameterization for per‑token gating in on‑policy knowledge distillation, unifying existing methods such as EOPD and ToDi as special cases. Experiments on TweetEval with a Qwen3 teacher‑student pair show that the full family of gating configurations outperforms single‑channel baselines in most cells, and dynamic gating beats static baselines in a majority of isolated comparisons. The authors present the framework mainly as a shared coordinate system for comparing gating designs rather than as definitive statistical evidence.

By Suwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang
arXiv Machine Learning
Jun 25

Why Do Accumulated Transformations Extrapolate?

arXiv:2606. 24975v1 Announce Type: new Abstract: PaTH Attention showed that replacing RoPE's position-indexed rotations with accumulated data-dependent Householder reflections yields strong length extrapolation, though performance degrades at extreme context lengths.

By Mahesh Godavarti