arXiv Machine Learning By Joshua Nunley

Kuramoto Attention: Synchronizing Self-Attention on the Torus

Read the original on arXiv Machine Learning →

arXiv:2606. 11585v1 Announce Type: new Abstract: We introduce Kuramoto attention, a self-attention layer in which each hidden coordinate is an angle.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 18

Attention as Frustrated Synchronization

arXiv:2606. 18694v1 Announce Type: new Abstract: A network of oscillators that synchronizes perfectly computes nothing further, so an attention architecture built from synchronization must locate its computation in structured departures from agreement.

By Joshua Nunley
arXiv Computation and Language
Sep 10

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance

The paper introduces a bounded spectral framework to analyze rotary attention in transformer language models, focusing on phase alignment, hidden‑state continuity, and semantic drift. It identifies ordered hidden‑state sequences as suitable domains for spectral decomposition, derives the Rotary Position Embedding (RoPE) attention score as a sum of magnitude‑weighted cosine terms, and proves a local stability lemma linking bounded phase displacement to pre‑softmax score degradation. By defining complex modal coordinates and a weighted coherence functional, the work distinguishes representational continuity from execution‑boundary admissibility, offering a theoretical program for when spectral structure explains continuity and when external governance is required.

By Abraham Chachamovits
arXiv AI
Jul 23

Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention

arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).

By Luis Rosario Freytes