arXiv Machine Learning By Andrew Gracyk

Observability conditions for neural state-space models with eigenvalues and their roots of unity

Read the original on arXiv Machine Learning →

The paper investigates observability in neural state‑space models, particularly the Mamba architecture, using tools from ordinary differential equations and control theory. It introduces several strategies—based on eigenvalues, roots of unity, permutations, Fourier transforms, and Vandermonde matrices—to enforce observability in high‑dimensional, learnable hidden states while maintaining computational efficiency. The authors also present a shared‑parameter construction for Mamba and a training algorithm that satisfies a Robbins‑Monro condition, contrasting it with classical procedures that fail to meet contraction requirements.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 11

The Spectral Neuron

arXiv:2608. 08003v1 Announce Type: cross Abstract: As machine learned models increase in complexity and expressive power, features of simpler models, such as interpretability and control over the shape of the modeled function are lost.

By Alex Shtoff
arXiv Machine Learning
Sep 7

The Sample Complexity of Learning Lipschitz Operators with respect to Gaussian Measures

The paper investigates how many linear samples are needed to learn Lipschitz operators under Gaussian measures. It establishes both lower and upper bounds on the Hermite polynomial approximation error and shows that the minimal worst‑case error cannot converge algebraically with the number of samples. However, if the covariance operator of the Gaussian measure decays rapidly, convergence rates arbitrarily close to any algebraic rate can be achieved.

By Ben Adcock, Michael Griebel, Gregor Maier
arXiv Machine Learning
2d ago

Same Loss, Different Gradients

arXiv:2609.38786v1 Announce Type: new Abstract: Differentiable learning typically assumes that the scalar objective evaluated in the forward pass and the gradient supplied to the optimizer in the bac...

By Ningkang Peng, Xiaoqian Peng, Yifan He, Anjie Hu, Chao Tan, Peirong Ma, Yanhui Gu