arXiv AI By Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu, Martin Benning, Gregory Slabaugh, Jiankang Deng, Ismail Elezi

MidSteer: Optimal Affine Framework for Steering Generative Models

Read the original on arXiv AI →

arXiv:2605. 05220v2 Announce Type: replace-cross Abstract: Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv Machine Learning
Jun 18

Concept Modulation Models: A Unified Framework for Identifiability and Extrapolation

arXiv:2606. 18509v1 Announce Type: new Abstract: Reliable generalization in conditional latent variable models requires understanding both identifiability and extrapolation: how observed variation across attributes determines latent structure, and how that structure determines distributions at unseen attributes.

By Soheun Yi, Yizhou Lu, Chandler Squires, Pradeep Ravikumar
arXiv Machine Learning
Jul 31

Dynamically Scaled Activation Steering

arXiv:2512. 03661v2 Announce Type: replace Abstract: Activation steering has emerged as a powerful method for guiding the behavior of generative models towards desired outcomes such as toxicity mitigation.

By Alex Ferrando, Xavier Suau, Jordi Gonz\`alez, Pau Rodriguez