arXiv:2607. 15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks.
By Felippe Alves, Renato Vicente
arXiv:2609.26067v1 Announce Type: new
Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions, increasing flexibility but also parameter memory be...
By Kazi Ahmed Asif Fuad, Lizhong Chen
arXiv:2608. 00859v1 Announce Type: new Abstract: Kolmogorov--Arnold Networks (KANs) replace scalar edge weights with learnable univariate functions parameterized by multiple basis coefficients.
By Kazi Ahmed Asif Fuad, Lizhong Chen
arXiv:2606. 02608v1 Announce Type: new Abstract: We study a Marchenko--Pastur (MP) random-matrix approach to pruning deep neural networks with very small post-pruning fine-tuning budgets.
By Leonid Berlyand, Theo Bourdais, Houman Owhad, Yitzchak Shmalo
FlashKAN introduces a new implementation of Kolmogorov‑Arnold Networks (KANs) that replaces the traditional Cox‑de Boor recursion with a truncated power form, allowing each uniform cubic B‑spline to be expressed as five shifted −(x)−^3 terms. The torch.compile‑fused implementation collapses these operations into a single GPU kernel, eliminating recursion, span lookup, and scatter‑gather steps. Additionally, the method includes a bounded‑coordinate stabilization to clamp inputs to [0, k+1], preventing catastrophic cancellation, and provides a production‑ready, open‑source package (pip install flashkan) as a drop‑in replacement for existing KAN layers.
By Naveen Mysore
arXiv:2608. 25807v1 Announce Type: new Abstract: Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central.
By K S Sesh Kumar
arXiv:2609. 01956v2 Announce Type: replace Abstract: Kolmogorov-Arnold Networks (KANs) place learnable B-spline activations on network edges rather than fixed activations on nodes.
By Naveen Mysore
arXiv:2607. 01478v1 Announce Type: cross Abstract: We measured quantization-induced decision-boundary changes using local logit-margin radii, first-order boundary displacement, normal variation, slice-boundary Jaccard distance, grid prediction changes, multiclass junction counts, and low-margin boundary-band flips.
By O. M. Kiselev
arXiv:2608. 00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture.
By Enzo Nicolas Spotorno, Josafat Leal Filho
The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.
By Ravi Satya Durga Prasad Yenugula
arXiv:2608. 07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head.
By Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau, Anass Belfatmi
Kolmogorov-Arnold Networks (KANs) replace fixed activations in deep architectures with learnable univariate edge functions, making the choice of edge parametrisation central. Existing variants rely on...