Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
arXiv:2609.01048v1 Announce Type: cross Abstract: Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as...
The paper investigates how a probe can decode in‑context bindings on model errors and how probe‑guided steering can repair them. It tracks probe accuracy, model output, and steering response across public pretraining and post‑training checkpoints, noting that probe accuracy improves during Pythia pretraining and that steering benefits grow with model size. The study also shows that decoders trained on final state or candidate logits do not outperform each other on late‑checkpoint errors, and presents an information‑theoretic counterexample explaining why decodability on errors alone cannot prove discarded output information.
The paper introduces Counteraction-Aware Multi-Teacher On-Policy Distillation (CaMOPD), a method designed to recover general capabilities in large language models while preserving domain-specific behavior. CaMOPD tackles two failure modes of standard Multi-Teacher On-Policy Distillation—conflicting recovery and preservation gradients, and weak correction signals—by using decoupled alternating training and selecting samples with large teacher‑student log‑probability gaps. Experiments on role‑play dialogue and medical reasoning QA show that CaMOPD outperforms baselines in general capability recovery while maintaining domain specialization, and gradient coherence analyses confirm more coherent correction signals.
The paper investigates how the upper spectral tails of weight matrices in decoder‑only transformer language models influence reasoning behavior. By performing controlled interventions on the query–key product and comparing them to factor‑level surgeries, the authors find that edits targeting the spectral tail more strongly affect model performance across multiple checkpoints and reasoning benchmarks. The study also explores how inverse participation ratios predict accuracy transitions and shows that tail‑aware low‑rank adaptations converge faster than standard methods.
arXiv:2609.36569v1 Announce Type: cross Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...