arXiv AI By Heyang Gong

From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics

Read the original on arXiv AI →

The paper introduces OAttention, a token‑level attention mechanism that assigns each token a presence coefficient based on its hidden representation. This coefficient both gates the token’s output and weights its contribution to other tokens, making zero‑vector tokens behave as true zeros and enabling exact null‑receiver, null‑source, and empty‑support properties. The authors extend this idea to local O‑components and an O‑Transformer, and demonstrate small performance changes when retrofitting a pretrained TabPFN model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 23

Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention

arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).

By Luis Rosario Freytes
arXiv Machine Learning
Aug 12

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

arXiv:2608. 10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws.

By Burc Gokden