arXiv AI By Xinrui Chen, Jianhao Zhang, Ou Wu, Di Gao

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

Read the original on arXiv AI →

arXiv:2606. 09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.