arXiv Machine Learning By Roman Maksimov, Vladimir Aletov, Vladimir Solodkin, Dmitry Bylinkin, Daniil Medyakov, Aleksandr Beznosikov

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

Read the original on arXiv Machine Learning →

arXiv:2608. 17836v1 Announce Type: new Abstract: As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.