arXiv Machine Learning By Sridhar Mahadevan

Kan Extension Transformers: A Categorical Unification of Attention, Diffusion, and Predict-Detach Self-Conditioning

Read the original on arXiv Machine Learning →

arXiv:2605. 27259v2 Announce Type: replace Abstract: We propose Kan Extension Transformers (KETs) as a categorical design language for a diverse group of Transformer implementations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

TANGO: Treating Tokens as Operators

The paper introduces TANGO, a Token‑Aggregated Nonlinear Gating Operator that blends cross‑token mixing and token‑wise transformation in transformer architectures. By computing a nonlinear gate per source token and averaging these gates for each destination, TANGO forms a source‑conditioned linear operator that improves predictive performance. Experiments on web text, formal mathematics, and code show that full‑prefix TANGO achieves the lowest test negative log‑likelihood across 16 settings, while a narrower variant offers substantial throughput gains with only a modest increase in loss.

By Joshua Nunley
arXiv AI
Jul 23

Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention

arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).

By Luis Rosario Freytes
arXiv Machine Learning
Jul 7

Legible-by-Construction: Attention and End-to-End Transformers

arXiv:2607. 04319v1 Announce Type: cross Abstract: A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a self-forgetting sequence quantifier - so its hidden units read as named logical operators at no cost to language-model quality.

By Mark Oskin