arXiv:2609.23900v1 Announce Type: new
Abstract: Tree speculative decoding verifies multiple candidate continuations in one target forward pass. For attention-only transformers, the verifier mainly ne...
By Zhiyuan Ma
TriPLU is a Trilinear Product Linear Unit that replaces the gated feed‑forward branch in tiny decoder‑only language models with a degree‑3 product‑only branch that multiplies three projected streams coordinate‑wise. In a character‑level TinyStories 1M‑byte prefix study, TriPLU achieves a mean best validation loss of 1.0637, outperforming closely matched SwiGLU (1.1017), a degree‑4 product control (1.0780), and a degree‑2 control (1.1026). In train‑only Byte‑BPE experiments, TriPLU also lowers validation and held‑out bits per byte on TinyStories and WikiText‑2 raw under low‑learning‑rate settings, with PMI‑slice evidence indicating gains on seen middle‑ and high‑PMI adjacent‑token pairs.
By He Zhang
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.
By Simon P. Villani
arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.
By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead.
"whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."
By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
TreeWY introduces a speculative verification method for Gated DeltaNet (GDN) hybrid models that eliminates the need for per-draft-state snapshots. By applying a tree‑structured WY transform to the gated delta rule, each draft node’s output is computed with a single triangular solve, and only the accepted state is reconstructed on commit. Benchmarks on Qwen3.5 35B and 397B show reduced memory pressure, higher throughput, and lower time‑to‑first‑token in memory‑bound scenarios, while enabling wider, higher‑acceptance draft trees.
By Sneha Murthy Ghantasala
arXiv:2605.21333v3 Announce Type: replace-cross
Abstract: Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combinati...
By Ting Liu
arXiv:2607. 27591v1 Announce Type: new Abstract: Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activation sparsification.
By Jinyi Liu, Wei Chen, Pengyu Chen, Xinyi Yuan, Minghe Bai, Guoquan Wu, Jun Wei
TinyCeNN-LM proposes a quality‑gated post‑training conversion framework that replaces attention in pretrained language models with CeNN‑inspired cellular‑recurrent layers. The method includes bounded local processing, compact recurrent memory, routing, fusion, and an accept‑or‑rollback validation step, with three specific implementations studied. Experiments on SmolLM2‑135M and Qwen3.5‑0.8B show that the conversion can accept certain layers while rejecting others based on representation fidelity and NLL thresholds, achieving minimal perplexity changes and modest downstream accuracy retention.
By Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya
The paper introduces a method for quantizing the fixed‑size recurrent states of hybrid language models to as few as four bits per token. By deriving distortion weights from the observability Gramian and combining them with normalized state ranges, the authors achieve mixed‑precision bit allocation without requiring calibration data, rotation, or additional training. The approach also logarithmically quantizes decay rates, yielding significant reductions in excess negative log‑likelihood—up to 27.9× better than seven baselines—while maintaining near‑FP32 performance at six bits.
By Hongren Chen, Jiayang He
The paper demonstrates that a byte‑level BPE tokenizer can be sliced to create multiple vocabulary sizes from a single trained model, preserving exact logits while reducing deployed weights by 66%. Experiments on 30 models show that while sliced models match the full model numerically, they underperform fixed‑cap specialists by a few percentage points in bits‑per‑byte. Multi‑cap training improves robustness to typographical noise, suggesting benefits from training across multiple granularities rather than from control tokens alone.
By Christos Koutsiaris