arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv Machine Learning
2d ago

Decodable In-Context State and Model Output Across Training

The paper investigates how a probe can decode in‑context bindings on model errors and how probe‑guided steering can repair them. It tracks probe accuracy, model output, and steering response across public pretraining and post‑training checkpoints, noting that probe accuracy improves during Pythia pretraining and that steering benefits grow with model size. The study also shows that decoders trained on final state or candidate logits do not outperform each other on late‑checkpoint errors, and presents an information‑theoretic counterexample explaining why decodability on errors alone cannot prove discarded output information.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
arXiv AI
Sep 15

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

The paper introduces Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD), a method designed to recover general capabilities in large language models while preserving domain-specific behavior. CaMOPD tackles two failure modes of standard Multi-Teacher On-Policy Distillation—conflicting recovery and preservation gradients, and weak correction signals—by using decoupled alternating training and selecting samples with large teacher‑student log‑probability gaps. Experiments on role‑play dialogue and medical reasoning QA show that CaMOPD outperforms baselines in general capability recovery while maintaining domain specialization, and gradient coherence analyses confirm more coherent correction signals.

By Tianlei Chen, Jiao Ou, Ziyuan Liu, Ruiming Tang, Jian Liang, Han Li
arXiv Machine Learning
Sep 23

Spectral Tail Interventions in Decoder-Only Language Models: Reasoning-Sensitive Weight Structure from Controlled Surgery

The paper investigates how the upper spectral tails of weight matrices in decoder‑only transformer language models influence reasoning behavior. By performing controlled interventions on the query–key product and comparing them to factor‑level surgeries, the authors find that edits targeting the spectral tail more strongly affect model performance across multiple checkpoints and reasoning benchmarks. The study also explores how inverse participation ratios predict accuracy transitions and shows that tail‑aware low‑rank adaptations converge faster than standard methods.

By Ibne Farabi Shihab, Sanjida Akhter, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma