arXiv AI By Fan L. Cheng, Nikolaus Kriegeskorte

Decomposing how prompting steers behavior

Read the original on arXiv AI →

arXiv:2606. 03093v1 Announce Type: new Abstract: Prompting steers large language models (LLMs) and vision-language models (VLMs) without weight updates, but it remains unclear how instruction changes reshape internal representations to produce behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

arXiv:2605.30117v2 Announce Type: replace Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VL...

By Haoyuan Shi, Xiancong Ren, Yingji Zhang, Qinfan Zhang, Jiayu Hu, Haozhe Shan, Han Dong, Jinpeng Lu, Yinda Chen, Yi Zhang, Yong Dai, Xiaozhu Ju
arXiv Computation and Language
3d ago

MetaSteer: Context-Conditioned, nonlinear Steering via Attention-Projection Adaptation

MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.

By Mehdi Jafari, Hao Xue, Flora Salim
arXiv AI
Aug 19

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

The paper audits frozen decoder‑only large language models (LLMs) on geometric reasoning tasks using parametric CAD constraints. It probes hidden states for linear decodability, forced‑choice generation, activation‑level influence, and behavioral steerability, finding that pretraining improves decoding of local geometric relations but not sketch‑level DOF status. The study shows that decodable information is not always actionable: generation often fails to express it, and steering interventions do not reliably control outputs, revealing divergences among decodability, generation, activation influence, and steerability.

By Man Liang, Xinzhao Cheng, Faizan Wajid