arXiv Machine Learning
2d ago

Decodable In-Context State and Model Output Across Training

The paper investigates how a probe can decode in‑context bindings on model errors and how probe‑guided steering can repair them. It tracks probe accuracy, model output, and steering response across public pretraining and post‑training checkpoints, noting that probe accuracy improves during Pythia pretraining and that steering benefits grow with model size. The study also shows that decoders trained on final state or candidate logits do not outperform each other on late‑checkpoint errors, and presents an information‑theoretic counterexample explaining why decodability on errors alone cannot prove discarded output information.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha