arXiv AI

Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science

arXiv:2607. 23634v1 Announce Type: cross Abstract: Attention enables context modeling via query-key scoring with softmax normalization.

arXiv Machine Learning
Sep 25

Beyond Pairwise Attention: Higher-Order Modular Attention for Efficient Sequence Learning

The paper introduces Higher-Order Modular Attention (HOMA), a new attention mechanism that combines standard pairwise self‑attention with an explicit triadic attention pathway. HOMA uses overlapping blocks, local windows, and a low‑rank projection to make triadic interactions tractable. Experiments on controlled PARITY and MATCH3 tasks, as well as TAPE benchmarks, show that HOMA matches or outperforms matched pairwise and purely triadic baselines, especially when dependencies extend beyond triadic order, and it often converges faster and uses parameters more efficiently.

By Shirin Amiraslani, Xin Gao
arXiv Machine Learning
Jun 9

Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization

arXiv:2510. 13554v2 Announce Type: replace-cross Abstract: The reasoning pattern of Large language models (LLMs) remains opaque, and reinforcement learning (RL) typically applies uniform credit across an entire generation, blurring the distinction between pivotal and routine steps.

By Yang Li, Zhichen Dong, Yuhan Sun, Weixun Wang, Shaopan Xiong, Yijia Luo, Jiashun Liu, Han Lu, Jiamang Wang, Wenbo Su, Bo Zheng, Junchi Yan
arXiv Machine Learning
Sep 25

Invertible Query-Key Coupling Composes with Attention Mechanisms

The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.

By Barak Gahtan, Alex M. Bronstein
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen
arXiv Machine Learning
Aug 27

A General-Purpose Framework for Chemical Reaction Representation with Atomic Correspondence and Flexible Condition Adaptation

The paper introduces Align-React, a chemical reaction representation learning framework that incorporates atomic correspondence between reactants and products, an adapter for embedding reaction conditions, and a Reaction-Center-Aware attention mechanism. These components enable the model to capture precise molecular transformations and focus on critical functional groups, leading to improved performance across a variety of organic reaction tasks. The framework outperforms existing architectures on most benchmark datasets.

By Kaipeng Zeng, Xianbin Liu, Yu Zhang, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv Statistics ML
6d ago

On the Expressive Power of Transformers for Contextual Relations

The paper investigates the theoretical expressive power of Transformers in modeling contextual relations. By framing a text as a distribution of representations and attention as a probabilistic relation, it connects attention normalization to optimal transport: softmax yields conditional relations, while Sinkhorn yields joint relations with fixed marginals. The authors prove universal approximation results, showing that Transformers with Sinkhorn normalization can represent any joint probability relation, whereas standard softmax Transformers can represent any conditional probability relation.

By Demi\'an Fraiman