arXiv Machine Learning By Mehrshad Saadatinia, Parsa Razmara, Ardalan Aryashad, Ali Abbasi, Seyedarmin Azizi

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

Read the original on arXiv Machine Learning →

arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 9

Distributed Sparse Interventions in Language Models

arXiv:2607. 07128v1 Announce Type: new Abstract: Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks.

By Maximilian S. Ernst (Max Planck School of Cognition, Center for Lifespan Psychology Max Planck Institute for Human Development, Machine Learning Group Technische Universit\"at Berlin), Lorenz Linhardt (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data), Aaron Peikert (Center for Lifespan Psychology Max Planck Institute for Human Development), Oliver Eberle (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data)