arXiv AI

A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

arXiv:2608. 13626v1 Announce Type: new Abstract: A hidden state signal can be decodable or causally usable without supporting a reusable action map.

arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
6d ago

Causal Retention in Interactive Agents: Interface Factorization and Selective Adaptation

The paper introduces the concept of causal retention in interactive agents, examining whether a frozen learned state can correctly answer a mechanism‑probe map that is fixed independently of training. It shows that for finite structural causal models the optimal probe error is a Bayes decision risk, vanishing only when each learning‑interface fiber lies within a single probe‑answer fiber, and provides theoretical results such as a posterior‑coverage theorem and an exact edit decomposition. Experiments on finite causal systems, continuous simulators, TD‑MPC2, and Qwen2.5‑7B‑Instruct demonstrate that causal retention can be achieved with high accuracy, outperforming task‑performance‑based approaches.

By Shengjun Zhang, Tingyi Liu, Dong Xie, Yunlong Dong, Xiang Wang, Cheng Zeng
arXiv AI
6d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv Machine Learning
Aug 5

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.

By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
arXiv Computation and Language
Aug 28

SCIT: Testing Causal Cache Carriers in Latent Chain-of-Thought Models

SCIT (Suffix Cache Interchange Test) is a causal protocol designed to identify which transformer components carry counterfactual computations in latent chain-of-thought models. By constructing exact source‑recipient counterfactuals and applying sufficiency tests, K/V splits, hidden‑state controls, and semantic source controls, SCIT demonstrates that counterfactual arithmetic primarily transfers through value‑cache suffix trajectories rather than hidden states or keys. The method reveals carrier‑regime shifts across different GPT‑2 checkpoints, providing a cache‑level diagnostic and a competence‑gated carrier map for arithmetic mechanisms.

By Yi Ding, Lijun Huang, Menglin Yang
arXiv AI
Sep 12

Portable Semantics, Private Dialects: Reuse and Negative Transfer in Latent Communication Between Language-Model Cells

The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.

By Narcis Marincat
arXiv Machine Learning
Aug 4

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

arXiv:2608. 02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall.

By Jingxi Wei