arXiv Machine Learning

BiKAN: Restoring Collapsed Basis of Binary Kolmogorov--Arnold Networks

arXiv:2608. 01490v1 Announce Type: new Abstract: Binarizing a polynomial Kolmogorov--Arnold Network (KAN) not only changes parameter precision, but also alters the function space available to each layer.

arXiv Machine Learning
Sep 3

FlashKAN: B-Spline KANs via Truncated Power Form

FlashKAN introduces a new implementation of Kolmogorov‑Arnold Networks (KANs) that replaces the traditional Cox‑de Boor recursion with a truncated power form, allowing each uniform cubic B‑spline to be expressed as five shifted −(x)−^3 terms. The torch.compile‑fused implementation collapses these operations into a single GPU kernel, eliminating recursion, span lookup, and scatter‑gather steps. Additionally, the method includes a bounded‑coordinate stabilization to clamp inputs to [0, k+1], preventing catastrophic cancellation, and provides a production‑ready, open‑source package (pip install flashkan) as a drop‑in replacement for existing KAN layers.

By Naveen Mysore
arXiv Machine Learning
Aug 4

An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models

arXiv:2608. 00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture.

By Enzo Nicolas Spotorno, Josafat Leal Filho
arXiv AI
Sep 25

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula