arXiv AI By Seyed Arshan Dalili, Ajay Narayanan Sridhar, Vijaykrishnan Narayanan, Mehrdad Mahdavi

Conditional Optimal Bridge for Riemannian Activation Steering

Read the original on arXiv AI →

arXiv:2607. 10517v1 Announce Type: cross Abstract: Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models

GeoSteer introduces a geometry-aware, optimization-based approach to norm-preserving activation steering in large language models. By formulating steering as a Riemannian optimization problem, it updates activations through a sequence of small geodesic steps guided by a learned nonlinear objective, avoiding fixed steering directions. Experiments on TruthfulQA, RealToxicityPrompts, and UltraFeedback show that GeoSteer consistently outperforms existing activation steering baselines, offering smoother, more stable, and more consistent steering behavior.

By Xuan Cuong Ngo, Hao Vo, Ngan Le
arXiv Machine Learning
Aug 19

FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers

FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.

By Sihan Wang, Jiayi Zhao, Qingyan Cao, Hongbo Yao, Lin Shu
arXiv Machine Learning
Sep 3

IDEEA: training-free Input-Dependent stEEring via Activation cluster matching

IDEA is a training‑free, input‑dependent steering method for large language models that matches activations to cluster‑specific directions aligned with a target concept. It clusters positive and negative activation supports per attention head, solves an optimal‑matching problem to create a pool of cluster‑conditional directions, and selects the best match for each input at inference time. This approach preserves the input’s original representation while improving the truth × info rate on TruthfulQA by an average of 9.9% (up to 23.5%) over input‑independent baselines.

By Zheng Wang, Muchen Li, Renjie Liao, Yan Leng
arXiv AI
Jul 23

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.

By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
arXiv Machine Learning
Jun 18

Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation

arXiv:2606. 18844v1 Announce Type: new Abstract: Self-distillation improves reasoning in large language models by using the model's own rollouts as training signal, typically through implicit logit-level alignment that minimizes KL divergence toward a privileged target distribution.

By Zhilin Huang, Hang Gao, Ziqiang Dong, Yuan Chen, Yifeng Luo, Chujun Qin, Jingyi Wang, Yang Yang, Guanjun Jiang