arXiv AI By Kyungmin Park, Taesup Kim

Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories

Read the original on arXiv AI →

arXiv:2606. 04778v1 Announce Type: new Abstract: Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.