arXiv Machine Learning

Sparsely gated tiny linear experts

arXiv:2606. 07414v1 Announce Type: new Abstract: Sparsity allows scaling model parameters without proportionally increasing computational cost.

arXiv Machine Learning
Aug 27

The Von-Neumann State-Space Transformer for neural decoding

The paper introduces the Von‑Neumann State‑Space Transformer (VN‑SST), a memory‑augmented Transformer that replaces the standard feed‑forward block with a low‑rank instruction bank. By decoding token‑specific operators from a low‑dimensional state‑space memory, VN‑SST achieves higher data‑efficiency and parameter‑efficiency on motor‑cortex neural‑decoding tasks and on small language‑model benchmarks. The model demonstrates that a compact instruction set can act as a control channel, improving performance without increasing accuracy through larger parameter counts.

By Morteza Sarafyazd