SIMS: Scale-Invariant Merit-Function-Based Scalarization for Multi-Task Learning proposes a new scalarization method for multi-task learning that is invariant to the relative scales of task losses. By using a logarithmic transformation, SIMS converts the multi-objective problem into a single objective that preserves weak Pareto optimality and allows a smooth surrogate with controllable approximation error. Experiments on standard multi-task benchmarks show that SIMS consistently outperforms existing scalarization methods and achieves state‑of‑the‑art performance.
By Zebin Chen, Fei Xing, Yang Chen, Hua Liu, Andy HF Chow, Yuhua Qian, Yu Zhang
arXiv:2604.13287v2 Announce Type: replace
Abstract: Weight pruning is a common technique for compressing large neural networks. We focus on the challenging post-training one-shot setting, where a pre...
By Gabriel Afriat, Xiang Meng, Shibal Ibrahim, Hussein Hazimeh, Rahul Mazumder
arXiv:2406. 09770v2 Announce Type: replace-cross Abstract: Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landscape and the expensive computational cost of training and evaluating models.
By Anke Tang, Li Shen, Yong Luo, Shiwei Liu, Han Hu, Bo Du, Dacheng Tao
HUANet is a deep neural network architecture that unrolls the Alternating Direction Method of Multipliers (ADMM) into a trainable model for accelerating parametric constrained convex optimization. It embeds a hard‑constrained neural network in each ADMM iteration, using a differentiable correction stage to enforce affine equalities of the primal subproblem. The method also incorporates first‑order optimality conditions into a self‑supervised training loss, and numerical experiments on benchmark problems and a control application demonstrate its effectiveness in speeding up constrained convex optimization.
By Trinh Tran, Binh Nguyen, Truong X. Nghiem
arXiv:2608. 09523v1 Announce Type: new Abstract: Deep neural network (DNN) training with stochastic gradient descent (SGD) and its variants achieves strong empirical performance, yet classical optimization theory does not fully explain this success.
By Binchuan Qi
arXiv:2602. 07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives.
By Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma
arXiv:2607. 06151v1 Announce Type: new Abstract: Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitting and reduced performance on unseen data.
By Yao Fu, Chunxia Zhang, Junmin Liu, Yihang Jin, Haishan Ye, Yuanao Yang
arXiv:2601. 16884v3 Announce Type: replace Abstract: We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks.
By Shijun Zhang, Zuowei Shen, Yuesheng Xu
arXiv:2606. 30813v1 Announce Type: cross Abstract: Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training.
By Haoming Meng, Anton Sugolov, Vardan Papyan
arXiv:2605. 22876v2 Announce Type: replace Abstract: Existing neural solvers for Multi-Objective Combinatorial Optimization Problems (MOCOPs) commonly adopt decomposition-based strategies that scalarize a MOCOP into multiple subproblems associated with distinct weight vectors.
By Xuan Wu, Jinbiao Chen, Yang Li, Lijie Wen, Chunguo Wu, Yuanshu Li, Yubin Xiao, Chunyan Miao, You Zhou, Di Wang
arXiv:2605. 18629v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features.
By Micha{\l} Brzozowski, Neo Christopher Chung
arXiv:2606. 02134v1 Announce Type: cross Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations.
By Konstantin Kaulen, Hadar Shavit, Holger H. Hoos