arXiv Machine Learning By Hyunjin Cho, Youngji Roh, Jaehyung Kim

Shared Semantics, Divergent Mechanisms: Unsupervised Feature Discovery by Aligning Semantics and Mechanisms

Read the original on arXiv Machine Learning →

arXiv:2606. 08236v1 Announce Type: cross Abstract: As large language models are increasingly deployed in high-stakes settings, there is a growing need for tools that audit not only model outputs but also the internal computations that produce them.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.