arXiv Machine Learning

The Automaton Underneath: The Additive Input Pathway Is a Parasitic Attractor for State Tracking in Householder Linear RNN

arXiv:2609. 18966v1 Announce Type: new Abstract: Linear RNNs with input-dependent Householder-product transitions (DeltaNet/DeltaProduct-class) can provably represent hard state-tracking automata, yet trained models fail to length-generalize -- a gap recent work attributes to optimization, without a causal account.

arXiv AI
6d ago

PolicyAttention: Softmax Attention Implements Policy Mirror Descent for Closed-Loop Control

The paper investigates whether causal softmax attention can realize policy mirror descent (PMD) as a repeated controller rather than a one‑step algebraic identity. It constructs a fixed causal‑softmax actor–environment–one‑step‑critic protocol, detailing actor, routing, sampling, and normalization residuals, and shows that a frozen one‑step audit model closely approximates PMD. Empirical results demonstrate that the learned actor with an exact one‑step critic achieves median policy loss only about 5% higher than the exact PMD oracle across multiple control settings.

By Yuhe Sui, Yingzhi Tang, Shufang Chen
arXiv AI
Aug 25

Improving Few-Step Language Flows with Untied Self-Conditioning

The paper introduces Untied Self-Conditioning, a sampler that corrects a train–inference mismatch in flow‑matching language models. By dampening redundant directions in the self‑conditioning input and approximating a step‑average prediction from history, the method improves generation quality without retraining. On LangFlow and ELF‑B datasets, it dramatically lowers perplexity and is preferred in the majority of pairwise comparisons.

By Bocheng Li, Linli Xu
arXiv AI
Jul 21

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.

By Irina Piontkovskaia, Sergey Nikolenko