An Exploratory Ablation of a Small MLA--SSM Hybrid Language Model
Read the original on arXiv AI →The paper reports a single‑seed ablation study of the TALH language model, which combines a Multi‑head Latent Attention (MLA) branch with a custom recurrent state‑space model (SSM). Five variants ranging from 117 M to 217 M active parameters per token were trained on a FineWeb sample, and the results show that removing the SSM branch causes the largest drop in validation perplexity (315) compared to removing MLA (239). A dense‑FFN hybrid achieved a perplexity of 231, outperforming the tested top‑2 ternary‑MoE hybrid (240) while using 3.87 GB less peak training memory, and MLA‑only exhibited the flattest time‑to‑first‑token curve on an Apple M3, though the dense Transformer was faster overall.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.