arXiv Machine Learning

The Conflict Between Logic and Memory: Learning Higher-Order Interactions in Shallow MLPs

The paper investigates how single‑hidden‑layer MLPs can fit training data yet fail to recover the underlying rule, focusing on higher‑order interactions and nuisance inputs. Using synthetic parity tasks, the authors benchmark different optimizers (SGD, Adam, Muon) and show that while all achieve perfect accuracy on second‑order interactions, performance drops sharply for higher orders, with Muon outperforming the others at fourth order. Experiments also reveal that freezing or removing nuisance‑related weights dramatically alters training outcomes, highlighting the role of nuisance learning in shaping the rules a shallow network can represent.

arXiv AI
3d ago

Unmerge: Efficient Machine Unlearning via Task Arithmetic

The paper introduces Unmerge, an efficient machine unlearning algorithm that treats unlearning as the inverse of task arithmetic. By representing the forget component as a low‑rank basis at each layer, Unmerge optimizes three goals—matching the merged vector, suppressing leakage, and bounding correction size—to limit forget leakage and retain damage. Experiments on ResNet‑50, ViT‑S/16, and Llama‑3.2‑3B show significant performance gains over existing methods while maintaining privacy and feature‑distribution fidelity.

By Haoran Tang, Andrew Tan, Rajiv Khanna
arXiv Machine Learning
Sep 1

PRIME: Mitigating Subgroup Optimization Competition in Shared CTR Top Networks with Plug-in Residual Input-Conditioned Mixture of Expert

PRIME is a plug‑in residual input‑conditioned mixture of experts that preserves the original dense prediction path while adding low‑rank, input‑dependent logit corrections. By initializing residuals to zero, PRIME matches the baseline dense model at training start and stabilizes conditional estimation with multi‑bag aggregation and EMA load biases. Experiments on Avazu and Criteo across 13 CTR architectures show modest AUC and LogLoss gains, with PRIME outperforming APG on FiBiNET and DCNv2 while using fewer parameters and lower latency.

By Heng Yao, Siyun Hou, Tianying Liu, Yulou Shu, Yong He, Chuan Yuan, Kaibin Qiu, Guowei Chen, Jiayu Zhao, Chao Yu, Ke Ding
arXiv AI
Jul 22

Soft-TransFormers for Continual Learning

arXiv:2411. 16073v4 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework that adapts a frozen pre-trained Transformer through task-specific soft subnetworks: real-valued multiplicative masks over the query, key, value, and output projections of selected self-attention layers.

By Haeyong Kang, Chang D. Yoo
arXiv AI
Sep 15

Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization

The paper introduces MACCHIATO, a training algorithm that builds a ReLU‑MLP from partial truth‑table data while simultaneously constructing an explicit Boolean circuit over AND, OR, and XOR gates that certifies the network’s computation. The method iteratively projects residuals onto low‑dimensional Boolean classes, compiles the resulting circuit into a ReLU‑MLP, and uses logic minimization and influence‑based variable selection to achieve a six‑layer network with provable truth‑table error bounds. Experiments on synthetic random‑junta tasks show that these certified networks outperform Adam‑trained MLPs in data‑sparse or projection‑aligned regimes and complete faster than flat ESPRESSO in certain settings.

By Hrad Ghoukasian, Anastasis Kratsios
arXiv AI
Sep 10

Equivariance Breaks the Learning Rate

arXiv:2609.08381v1 Announce Type: cross Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...

By Andrei Manolache, Mathias Niepert
arXiv Machine Learning
Jun 9

Convergence Bound and Critical Batch Size of Muon Optimizer

arXiv:2507. 01598v5 Announce Type: replace Abstract: Muon, a recently proposed optimizer that leverages the inherent matrix structure of neural network parameters, has demonstrated strong empirical performance, indicating its potential as a successor to standard optimizers such as AdamW.

By Naoki Sato, Hiroki Naganuma, Hideaki Iiduka