← Back to all news
arXiv AI August 25, 2026 By Marek Mateusz Kowalski, Joshua Fonseca Rivera, Uzay Macar, David Demitri Africa

Measuring Activation Control in Large Language Models

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • diffusion
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jul 29

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.

By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
llmssafety
More like this →
arXiv Machine Learning
Jun 9

Beyond Linear Activation Steering: Invertible Latent Transformations for Controlling LLM Behavior

arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.

By Tuc Nguyen, Thai Le
llmsdiffusionbenchmarkssafety
More like this →
arXiv AI
Jul 13

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs

arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.

By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
llmsfine-tuningmultimodalsafety
More like this →
arXiv AI
Jun 9

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).

By Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu
llmsfine-tuningsafety
More like this →
arXiv AI
Jun 12

Prefill Awareness in Large Language Models

arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.

By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
llmsagentsbenchmarkssafety
More like this →
arXiv AI
2d ago

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...

By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
llmsfine-tuningbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea