arXiv AI

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.

arXiv Machine Learning
Sep 22

Testing the Construct Validity of a Functional Valence Axis in LLM Agents

The paper investigates whether contrastive activation directions in large language models truly capture a functional valence axis or merely reflect correlated features of the contrast used to extract them. By conducting controlled interventions in a maze task across multiple LLM checkpoints, the authors show that directions fitted on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer degrades, indicating the direction remains strongly conditioned on prior announcements and does not represent a history‑invariant scalar valence state.

By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng
arXiv Computation and Language
Sep 21

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.

By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao
arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
Aug 6

Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation

arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.

By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng