arXiv Machine Learning By Pankayaraj Pathmanathan, Furong Huang

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

Read the original on arXiv Machine Learning →

arXiv:2607. 27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
4d ago

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

arXiv:2608.03275v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve...

By Yiming Zeng, Lei Lu, Zexin Li, Zhuochun Li, Dehai Min, Shuoqiu Li, Shuyi Liao, Xidong Wu, Zeyu Zhang, Minmei Wang, Yu Zhao, Tingting Yu, Shangqian Gao