arXiv Machine Learning By Zirui Yan, Dennis Wei, Dmitriy A. Katz, Prasanna Sattigeri, Ali Tajer

Multi-component Causal Tracing in Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 03085v1 Announce Type: new Abstract: Causal tracing systematically intervenes on a large language model's (LLM's) internal representations to uncover and quantify the causal pathways linking specific inputs or computations to specific metrics of interest, quantifying the LLM's behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.