arXiv:2404.07729v2 Announce Type: replace
Abstract: Continual learning (CL) evaluates adaptability in learning solutions to retain knowledge. Our research addresses the challenge of catastrophic forg...
By Nadia Nasri, Carlos Guti\'errez-\'Alvarez, Sergio Lafuente-Arroyo, Saturnino Maldonado-Basc\'on, Roberto J. L\'opez-Sastre
CoMerge is a conflict‑driven preference optimization framework for merging multiple expert language models into a single multi‑task model without full retraining. It treats model merging as a preference optimization problem, using self‑supervised, conflict‑driven hard negative samples derived from naive merging defects to refine lightweight, tensor‑wise merging coefficients. Experiments show CoMerge achieves near‑perfect performance on MergeBench and improves instruction‑following and safety on Llama‑3.1‑8B‑Instruct while optimizing only 1,445 scalar coefficients.
By Mingjie Zheng, Zihao Chen, Wenqing Chen, Weile Yuan, Zhixuan Chu, Jianxing Yu, Zibin Zheng
CoMerge introduces a conflict‑driven preference optimization framework for merging multi‑task large language models, reframing merging as a preference problem that uses self‑supervised hard negative samples derived from naive merging defects. By optimizing lightweight, tensor‑wise merging coefficients, the method mitigates parameter‑space conflicts while preserving task‑specific capabilities. Experiments show CoMerge achieves an average normalized performance of 0.9968 on MergeBench and improves conflict‑sensitive tasks on Llama‑3.1‑8B‑Instruct, outperforming both data‑free and data‑driven baselines while optimizing only 1,445 scalar coefficients.
arXiv:2608.31096v1 Announce Type: cross
Abstract: Class-incremental learning (CIL) requires a model to incrementally learn tasks that contain new classes without accessing earlier training data while...
By Yunxiang Fu, Meng Lou, Yizhou Yu
arXiv:2504. 13822v3 Announce Type: replace-cross Abstract: The emergence of large pre-trained networks has revolutionized the AI field, unlocking new possibilities and achieving unprecedented performance.
By Eric Nuertey Coleman, Luigi Quarantiello, Ziyue Liu, Qinwen Yang, Samrat Mukherjee, Julio Hurtado, Vincenzo Lomonaco
arXiv:2603. 12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting, a critical limitation of the static pre-training paradigm inherent to modern LLMs.
By Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin