arXiv Machine Learning

Invertible Query-Key Coupling Composes with Attention Mechanisms

The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.

arXiv AI
Jul 23

Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention

arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).

By Luis Rosario Freytes
arXiv Computer Vision
Aug 31

Spectral Query-Key Product Weight Steering for Training-Free VLM Hallucination Mitigation

The paper introduces QK Product Steering, a data‑free, training‑free method that edits the query‑key product in vision‑language models to reduce object hallucination. By suppressing a few dominant singular modes in selected middle layers and mapping the edited product back to query weights, the approach lowers hallucination rates without affecting inference cost. Experiments on three GQA‑based VLMs show a 4.0% average reduction in CHAIR$_s$, with the effect localized to symmetric mutual‑attention channels.

By Karn Tiwari, Varnith Chordia, Prathosh A P