arXiv AI By Sida Liu, Feijiang Han

ICA Lens: Interpreting Language Models Without Training Another Dictionary

Read the original on arXiv AI →

arXiv:2606. 11722v1 Announce Type: cross Abstract: Finding interpretable directions in language-model representations is critical for understanding and controlling model behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.