arXiv Machine Learning By Yuting Ge, Pengju Yang, Mingkai Nie

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

Read the original on arXiv Machine Learning →

arXiv:2608. 07323v1 Announce Type: new Abstract: We test whether decoder-only language-model FFNs require SwiGLU's open positive tail.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 24

TriPLU: Bypassing the Gate with Direct Trilinear Product FFNs in Tiny Language Models

TriPLU is a Trilinear Product Linear Unit that replaces the gated feed‑forward branch in tiny decoder‑only language models with a degree‑3 product‑only branch that multiplies three projected streams coordinate‑wise. In a character‑level TinyStories 1M‑byte prefix study, TriPLU achieves a mean best validation loss of 1.0637, outperforming closely matched SwiGLU (1.1017), a degree‑4 product control (1.0780), and a degree‑2 control (1.1026). In train‑only Byte‑BPE experiments, TriPLU also lowers validation and held‑out bits per byte on TinyStories and WikiText‑2 raw under low‑learning‑rate settings, with PMI‑slice evidence indicating gains on seen middle‑ and high‑PMI adjacent‑token pairs.

By He Zhang
arXiv AI
Aug 20

Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining

The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.

By Godwin Abuh Faruna
arXiv Computation and Language
Sep 23

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.

By Simon P. Villani
arXiv Machine Learning
Aug 4

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.

By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li