arXiv Computation and Language

Phase Structure in Rotary Attention: A Spectral Framework for Semantic Continuity and Execution-Boundary Governance

The paper introduces a bounded spectral framework to analyze rotary attention in transformer language models, focusing on phase alignment, hidden‑state continuity, and semantic drift. It identifies ordered hidden‑state sequences as suitable domains for spectral decomposition, derives the Rotary Position Embedding (RoPE) attention score as a sum of magnitude‑weighted cosine terms, and proves a local stability lemma linking bounded phase displacement to pre‑softmax score degradation. By defining complex modal coordinates and a weighted coherence functional, the work distinguishes representational continuity from execution‑boundary admissibility, offering a theoretical program for when spectral structure explains continuity and when external governance is required.

arXiv AI
1d ago

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

The paper introduces Time‑Frequency Geometric Cross‑Attention (TFGCA), a module that enhances vision‑language‑action models by decomposing action chunks into time‑frequency tokens using a learnable wavelet transform. TFGCA fuses dot‑product similarity with wedge‑product magnitude to better capture both frequency‑based smooth trends and cross‑phase orthogonal motion structures. When added to a pretrained VLA model, it yields significant performance gains across in‑distribution and out‑of‑distribution benchmarks, including a 28.5‑point improvement under RoboTwin domain randomization and an 11.67‑point increase on real‑robot AgiBot A2 tasks.

By Shengye Dong, Haochen Niu, Hao Liu, Peiwen Lin, Chuang Wang, Shanmin Pang
Hugging Face Trending Papers
Aug 18

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

The paper argues that language operates with two parameters: amplitude, which measures how often words co‑occur, and phase, a signed relational factor that determines how co‑activated meanings combine and can reverse a meaning’s contribution. Unlike amplitude, phase is not captured by standard word embeddings or transformer attention weights and is indexed to individuals and dyadic interactions. The authors propose six empirical predictions to test phase’s role and suggest that future language models should incorporate agent‑indexed, phase‑bearing semantic states.

arXiv AI
Sep 3

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

The paper proposes an encoder Transformer that explicitly separates semantic, absolute positional (AP), and relative positional (RP) information, restricting the masked‑language‑modeling objective to the semantic stream. This disentanglement reveals that the AP subspace collapses into a low‑frequency two‑dimensional manifold reflecting document structure, that attention heads specialize into structure‑ and semantic‑oriented groups with RP supporting only the latter, and that standard positional encodings fail to robustly encode macroscopic structure. The approach preserves positional encoding and improves performance on 49 out of 65 linguistic phenomena in the Flash‑Holmes probing benchmark.

By Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski
arXiv Machine Learning
Jun 24

Selective Rotary Position Embedding

arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.

By Sajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter, Antonio Orvieto, Volkan Cevher
arXiv AI
Sep 1

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

arXiv:2605.30117v2 Announce Type: replace Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VL...

By Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju