arXiv:2609.38356v1 Announce Type: new
Abstract: Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continu...
By Sima Hashemi, Daniel Durstewitz, Georgia Koppe
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
arXiv:2608. 04060v1 Announce Type: cross Abstract: Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transition models are typically opaque neural maps.
By Yongchao Huang
arXiv:2607. 08234v1 Announce Type: cross Abstract: Real-world time series exhibit complex dynamics characterized by multiple simultaneous temporal patterns: short-term fluctuations, periodic seasonal cycles, long-term trends, and irregular abrupt changes.
By Sumit Satishrao Shevtekar, Chandresh Kumar Maurya
arXiv:2602. 15649v2 Announce Type: replace Abstract: In dynamical systems reconstruction (DSR) we aim to recover the dynamical system (DS) underlying observed time series.
By Alena Br\"andle, Lukas Eisenmann, Florian G\"otz, Daniel Durstewitz
arXiv:2606. 18457v1 Announce Type: new Abstract: Recurrent networks can contain substantial functional redundancy in weight space: changing a recurrent matrix may leave the input-output rollout nearly unchanged on a task distribution, while similar-scale changes can destroy the same behavior.
By Simon Dr\"ager
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
By Pedro Jim\'enez-Gonz\'alez, Miguel C. Soriano, Lucas Lacasa
The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.
By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
arXiv:2607. 04993v1 Announce Type: cross Abstract: Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them.
By Thomas Hofmann
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.
By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici