arXiv AI By Yixin Tan, Zhe Yu, Rui Wen, Jun Sakuma

One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs

Read the original on arXiv AI →

arXiv:2512. 14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.