arXiv Machine Learning By Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey

Mechanisms of Introspective Awareness

Read the original on arXiv Machine Learning →

arXiv:2603. 21396v5 Announce Type: replace Abstract: Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.