arXiv Machine Learning

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

arXiv:2608. 07323v1 Announce Type: new Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail.

arXiv Machine Learning
Aug 24

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

TriPLU is a Trilinear Product Linear Unit that replaces the gated feed‑forward branch in tiny decoder‑only language models with a degree‑3 product‑only branch that multiplies three projected streams coordinate‑wise. In a character‑level TinyStories 1M‑byte prefix study, TriPLU achieves a mean best validation loss of 1.0637, outperforming closely matched SwiGLU (1.1017), a degree‑4 product control (1.0780), and a degree‑2 control (1.1026). In train‑only Byte‑BPE experiments, TriPLU also lowers validation and held‑out bits per byte on TinyStories and WikiText‑2 raw under low‑learning‑rate settings, with PMI‑slice evidence indicating gains on seen middle‑ and high‑PMI adjacent‑token pairs.

By He Zhang
arXiv AI
Aug 20

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.

By Godwin Abuh Faruna
arXiv Computation and Language
Sep 23

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

By Simon P. Villani
arXiv Machine Learning
Aug 4

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.

By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
arXiv AI
Aug 24

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

TreeWY introduces a speculative verification method for Gated DeltaNet (GDN) hybrid models that eliminates the need for per-draft-state snapshots. By applying a tree‑structured WY transform to the gated delta rule, each draft node’s output is computed with a single triangular solve, and only the accepted state is reconstructed on commit. Benchmarks on Qwen3.5 35B and 397B show reduced memory pressure, higher throughput, and lower time‑to‑first‑token in memory‑bound scenarios, while enabling wider, higher‑acceptance draft trees.

By Sneha Murthy Ghantasala
arXiv AI
Sep 21

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

TinyCeNN-LM proposes a quality‑gated post‑training conversion framework that replaces attention in pretrained language models with CeNN‑inspired cellular‑recurrent layers. The method includes bounded local processing, compact recurrent memory, routing, fusion, and an accept‑or‑rollback validation step, with three specific implementations studied. Experiments on SmolLM2‑135M and Qwen3.5‑0.8B show that the conversion can accept certain layers while rejecting others based on representation fidelity and NLL thresholds, achieving minimal perplexity changes and modest downstream accuracy retention.

By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
arXiv Machine Learning
Sep 29

Low-Bit Recurrent States in Hybrid Language Models

The paper introduces a method for quantizing the fixed‑size recurrent states of hybrid language models to as few as four bits per token. By deriving distortion weights from the observability Gramian and combining them with normalized state ranges, the authors achieve mixed‑precision bit allocation without requiring calibration data, rotation, or additional training. The approach also logarithmically quantizes decay rates, yielding significant reductions in excess negative log‑likelihood—up to 27.9× better than seven baselines—while maintaining near‑FP32 performance at six bits.

By Hongren Chen, Jiayang He
arXiv Computation and Language
Aug 31

Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result

The paper demonstrates that a byte‑level BPE tokenizer can be sliced to create multiple vocabulary sizes from a single trained model, preserving exact logits while reducing deployed weights by 66%. Experiments on 30 models show that while sliced models match the full model numerically, they underperform fixed‑cap specialists by a few percentage points in bits‑per‑byte. Multi‑cap training improves robustness to typographical noise, suggesting benefits from training across multiple granularities rather than from control tokens alone.

By Christos Koutsiaris