arXiv AI

Mahalanobis-Based Multi-Head Attention for Complex State Propagation

The paper introduces Mahalanobis-Based Multi-Head Attention for Complex State Propagation (MHA‑CSP), a new attention mechanism that replaces the standard dot‑product with a Mahalanobis distance‑based RBF kernel. This approach enables infinite‑dimensional feature space attention without extra parameters, allows direct construction of Tree Attention via LogSumExp correction, and incorporates an attention meshing mechanism for cross‑head collaboration. Experiments show that with only 119K parameters and teacher forcing applied only at the final hidden state, MHA‑CSP outperforms Transformer and GCN baselines on long‑sequence state tracking tasks, demonstrating efficient structured reasoning.

arXiv AI
Sep 11

Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models

The paper introduces Time‑Frequency Geometric Cross‑Attention (TFGCA), a module that enhances vision‑language‑action models by decomposing action chunks into time‑frequency tokens using a learnable wavelet transform. TFGCA fuses dot‑product similarity with wedge‑product magnitude to better capture both frequency‑based smooth trends and cross‑phase orthogonal motion structures. When added to a pretrained VLA model, it yields significant performance gains across in‑distribution and out‑of‑distribution benchmarks, including a 28.5‑point improvement under RoboTwin domain randomization and an 11.67‑point increase on real‑robot AgiBot A2 tasks.

By Shengye Dong, Haochen Niu, Hao Liu, Peiwen Lin, Chuang Wang, Shanmin Pang
arXiv Computer Vision
Aug 27

Not All Attention Heads Contribute to Critical Visual Token Selection: Head-Aware Pruning Matters More

The paper shows that only a small subset of attention heads in vision-language models is responsible for selecting critical visual tokens. By pruning tokens based on similarity before LLM reasoning and then applying head‑aware pruning during reasoning, the proposed ProViP framework achieves high task performance with significant speedups. Experiments on LLaVA‑1.5‑7B demonstrate 95.9% performance retention and a 1.62× inference speedup at an 88.9% pruning ratio.

By Chaofang Ma, Lin Jiang, Carol Jingyi Li, Xingyu Liu, Zeyu Li, Jiang Xu, Wei Zhang
arXiv Machine Learning
Sep 18

Relational Attention for Data-Efficient Language Modeling

Relational BabyLM is a decoder‑only Transformer that replaces standard self‑attention with a Dual Attention Transformer (DAT) to separate object‑level lexical features from structural/relational information. The model incorporates a Next‑Latent Prediction objective to compress history into a dense belief state and introduces a RoPE‑based symbol‑retrieval mechanism. On the BabyLM 2026 challenge, the best model ranks 6th overall and 3rd on the NLP‑task subset, outperforming GPT‑2 on most benchmarks and achieving the highest EWoK score among strict‑track entries.

By Adrian Brasoveanu, Ece Takmaz, Jakub Dotla\v{c}il