Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. We introduce the quadrilateral loss, a differentiable penalty that treats additivity as a measurable behavior instead: a second-order mixed difference on pairs of training points swapping one coordinate, which vanishes if and only if the coordinate carries no interaction, remains informative for piecewise-linear networks, and equals in expectation the per-coordinate interaction mass of the interventional Shapley-GAM.
arXiv:2606. 27759v1 Announce Type: new Abstract: Training binary neural networks (BNNs) from scratch is dominated by the straight-through estimator (STE), whose forward/backward mismatch produces severe accuracy degradation as networks deepen.
By Evan Gibson Smith, Bashima Islam
arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
By Ethan Smith
arXiv:2605.06240v2 Announce Type: replace-cross
Abstract: Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inheri...
By Amirhossein Yousefiramandi
The paper investigates how single‑hidden‑layer MLPs can fit training data yet fail to recover the underlying rule, focusing on higher‑order interactions and nuisance inputs. Using synthetic parity tasks, the authors benchmark different optimizers (SGD, Adam, Muon) and show that while all achieve perfect accuracy on second‑order interactions, performance drops sharply for higher orders, with Muon outperforming the others at fourth order. Experiments also reveal that freezing or removing nuisance‑related weights dramatically alters training outcomes, highlighting the role of nuisance learning in shaping the rules a shallow network can represent.
By Gongyue Zhang, Honghai Liu
arXiv:2607. 27255v1 Announce Type: cross Abstract: Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization.
By Yanli Yan, Yuanzheng Li, Yong Zhao, Hongbo Guo, Shoudong Han