arXiv:2602. 18948v2 Announce Type: replace Abstract: Transformer models contain substantial internal redundancy arising from coordinate-dependent representations and continuous symmetries, in model space and in head space, respectively.
By J. Fran\c{c}ois, L. Ravera
arXiv:2606. 04032v1 Announce Type: cross Abstract: Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role.
By Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
arXiv:2604. 15010v2 Announce Type: replace-cross Abstract: When do transformers commit to a decision, and what prevents them from correcting it?
By \'Eric Jacopin
Activation patching reveals how facts are stored, routed, and read out across transformer layers, and why the residual stream does most of the work The post A Three-Phase Factual Recall Circuit in Gemma-2B and Gemma-12B-IT appeared first on Towards Data Science .
By Subhanga Upadhyay
arXiv:2606. 17830v1 Announce Type: cross Abstract: Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through functional equivalence.
By Viet-Hoang Tran, Vinh Khanh Bui, Van-Hoan Trinh, Tan Lai Ngoc, Tan M. Nguyen
arXiv:2605. 30836v2 Announce Type: replace Abstract: Recent SVD based compression methods for large language models like SVD LLM and Basis Sharing can be unified under one optimization problem.
By Snigdha Chandan Khilar