Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization. This raises a reliability question: whether a model refitted on pooled data preserves an action ordering supported by both sources.
arXiv:2607. 11347v1 Announce Type: new Abstract: Neural networks increasingly guide decisions in high-stakes domains such as medical diagnosis, credit approval, and energy bidding.
By Manli Yan, Yuebin Lin, Yaowen Yu, Yong Zhao
arXiv:2604. 27733v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective.
By Mehryar Mohri, Yutao Zhong
arXiv:2602. 24266v2 Announce Type: replace-cross Abstract: Which internal mechanisms of a neural network can be replaced while preserving the computation it performs?
By Amir Asiaee
arXiv:2607. 20201v1 Announce Type: cross Abstract: Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally.
By Antonio Di Cecco
Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.
By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts