arXiv Machine Learning By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng

Testing the Construct Validity of a Functional Valence Axis in LLM Agents

Read the original on arXiv Machine Learning →

The paper investigates whether contrastive activation directions in large language models truly capture a functional valence axis or merely reflect correlated features of the contrast used to extract them. By conducting controlled interventions in a maze task across multiple LLM checkpoints, the authors show that directions fitted on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer degrades, indicating the direction remains strongly conditioned on prior announcements and does not represent a history‑invariant scalar valence state.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.

By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng
arXiv Computation and Language
Sep 21

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.

By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao