arXiv:2510. 15814v2 Announce Type: replace-cross Abstract: Universality results for equivariant neural networks remain rare.
By Marco Pacini, Mircea Petrache, Bruno Lepri, Shubhendu Trivedi, Robin Walters
arXiv:2606. 02490v1 Announce Type: new Abstract: This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose.
By Antonin Oswald, Estelle Massart
This work studies neural architectures for classifying symmetric positive-definite matrices, focusing on congruence-like layers, in which the input matrix is multiplied on the left and right by a (possibly rectangular) weight matrix $W$ and its transpose. Such layers lie at the core of the celebrated SPDNet and have also been employed independently for dimensionality reduction on positive-definite data.
The paper introduces the directional linear separability measure (D‑LSM) to quantify how much linear separability is preserved or improved by injective affine maps in neural networks. It characterizes the geometry supporting D‑LSM, proves its invariance under injective affine embeddings, and derives conditions for gated activations (ReLU, GELU, SiLU) to preserve and recover samples. Experiments validate the theoretical bounds, demonstrate affine‑tube constructions that achieve guaranteed recovery, and apply the method to Vision Transformer representations to obtain early post‑activation separability certificates.
By Yi Wei, Xuan Qi, Suorong Yang, Furao Shen
arXiv:2607. 07035v1 Announce Type: cross Abstract: The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks.
By Changcun Huang
arXiv:2606. 16028v1 Announce Type: new Abstract: Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models.
By Thomas Dittrich, Oliver Potocki, Philipp Grohs