arXiv Machine Learning By Leona Hioki

Complex-Valued Phase-Coherent Transformer

Read the original on arXiv Machine Learning →

The paper introduces the Phase-Coherent Transformer (PCT), a complex-valued architecture that replaces traditional softmax attention with a real-valued, smooth gate applied to L2-normalised query-key similarities. PCT eliminates token competition, preserving phase information across layers, and demonstrates strong generalisation on a variety of mid-scale benchmarks, outperforming both standard softmax Transformers and other complex-valued counterparts. Experiments confirm that the gate design is essential: preserving negatively aligned phase components is crucial for performance, while violating these conditions leads to degradation or collapse on long-range tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 25

Do Value Vectors in Deep Layers Need Context from the Residual Stream?

The paper investigates whether deep transformer layers require context from the residual stream to compute value vectors. It finds that allowing deeper layers to use a context‑free value vector—preserving original token information—significantly improves performance, and adding context afterward yields little extra benefit. The authors introduce Bank of Values (BoV), a lookup table of token‑specific value vectors for the last third of layers, which reduces compute and memory while matching or surpassing prior methods on large models.

By Muyu He, Yuchen Liu, Qingya Huang, Li Zhang