arXiv:2606. 14970v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) has become a central application of modern optimization, enabling pretrained models to adapt to diverse downstream tasks and domain-specific data.
By Dmitriy Bystrov, Daniil Medyakov, Dmitry Bylinkin, Aleksandr Beznosikov
arXiv:2606. 10989v1 Announce Type: new Abstract: Large language model unlearning aims to suppress designated undesirable knowledge while preserving benign capabilities.
By Bocheng Ju, Jianhua Wang, Chengliang Liu, Xiaolin Chang
Large language model unlearning aims to suppress designated undesirable knowledge while preserving benign capabilities. Many unlearning objectives focus on suppressing undesired answers, while recent target-guided variants specify replacement behavior but still leave update locality largely unconstrained.
arXiv:2601. 09361v4 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models.
By Jiaying Zhang, Lei Shi, Jiguo Li, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He
arXiv:2603.28678v2 Announce Type: replace
Abstract: We introduce PACE, a backpropagation-free continual test-time adaptation system that directly optimizes the affine parameters of normalization laye...
By Damian S\'ojka, Sebastian Cygert, Marc Masana
arXiv:2608. 01422v1 Announce Type: cross Abstract: Machine unlearning seeks to remove targeted information from trained models without requiring costly retraining.
By Tyler Lizzo, Larry Heck
The paper introduces EoupCT, a framework that estimates and orthogonalizes unknown pre‑training gradients to mitigate catastrophic forgetting during continual fine‑tuning of large language models. It generates pseudo data most susceptible to forgetting using a learnable soft prompt with Gumbel‑Softmax, then jointly optimizes model parameters and the prompt via a first‑order Pareto optimizer to enforce orthogonality between new task updates and the estimated gradients. Experiments on multiple LLMs show that EoupCT preserves both task‑specific performance and the models’ inherent general‑purpose knowledge.
By Bing Wang, Changchun Li, Xin-Qiang Cai, Lin Yuanbo Wu, Ximing Li, Gang Niu, Masashi Sugiyama
arXiv:2606. 18024v1 Announce Type: cross Abstract: Catastrophic forgetting in continual adaptation is usually studied through parameter drift, replay, or distillation, but these views do not identify which output-space directions are vulnerable.
By Ido Nitzan Hidekel, Dan Raviv
The paper introduces Foundation Preserving LoRA (FoLoRA), a forgetting‑aware optimization framework that balances adaptation to downstream tasks with preservation of pretraining capabilities. FoLoRA uses a first‑order preservation condition to define a forgetting penalty based on pretraining‑proxy activations and a task utility from downstream activations, scoring update directions via a generalized Rayleigh quotient. This spectral coordinate system enables gated Adam updates that reduce low‑utility, high‑penalty directions, and the method constructs pretraining proxy calibration data by sampling from the pretrained model. Experiments on math, code, and instruction‑following tasks demonstrate that FoLoRA achieves a stronger balance between target task performance and aggregate preservation of non‑target capabilities compared to baselines.
By Dongjun Kim, Adrian de Wynter, Huancheng Chen, Heasung Kim, Haris Vikalo
arXiv:2608. 11674v1 Announce Type: cross Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation.
By Kai Yang, Jingwei Xu, Wanyu Wang, Kai-Yuan Guo, Zhenbo Yu, Yi Wang, Yu Qiao
arXiv:2606. 00147v1 Announce Type: cross Abstract: Domain-specific supervised fine-tuning (SFT) often improves in-domain performance at the cost of degrading a model's general capabilities.
By Yuduo Li, Xiaofeng Shi, Qian Kou, Longbin Yu, Hua Zhou
arXiv:2610.00431v1 Announce Type: new
Abstract: Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new t...
By Hang Yin, Haozhe Wang, Yuhua Luo, Zhangqi Pan, Xiaoxing Wang, Junchi Yan