arXiv AI By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

Read the original on arXiv AI →

arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.