arXiv Machine Learning

Local Gains and Fixed-Assignment Set Losses in Shared Set Decoders

arXiv:2608. 14717v1 Announce Type: cross Abstract: A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it.

arXiv Machine Learning
Sep 11

Published Unlearning Numbers Move Per Checkpoint, and Not Because the Removed Data Survives: An Audit of 263 Released Batch-Normalized Checkpoints

The audit examines 263 batch‑normalized checkpoints released by arXiv, finding that refitting models on retained data at identical weights shifts 47 of 221 checkpoints beyond the spread indicated by their own release seeds. This movement is attributed to checkpoint properties rather than the survival of removed data, as swapping removed records for kept ones barely changes the published state. The study concludes that releases should specify the fitting convention used, especially for batch‑normalized vision models.

By Junlong Shen Xingyu Li
arXiv AI
Sep 24

Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?

The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.

By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng
arXiv AI
Sep 18

Correct Now, Insufficient Later: Auditing Update Sufficiency in Context Compression

The paper investigates how memory systems can answer a current query correctly yet fail to retain distinctions needed for later updates. Using a paired‑history audit, the authors evaluate 24 history pairs across six synthetic mechanisms and two model backends, achieving perfect reveal accuracy on DeepSeek and high accuracy on GLM. Record‑level audits reveal specific failures in structured reveal memories and frontier late‑reference adequacy, and the authors test a label‑equivariant repair that only partially restores correctness.

By Guangzhe Zhang