arXiv AI By Rui Wu, Tong Che

Reference Feature Atlases for Mechanistic Auditing of Language Models

Read the original on arXiv AI →

arXiv:2607. 22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.