arXiv:2606.14636v3 Announce Type: replace
Abstract: Many two-stage estimators assess the first-stage learner by prediction error, even when the next stage uses its residual. In control-function instr...
By Rui Wu, Zongyuan Chen, Hong Xie, Defu Lian, Enhong Chen
arXiv:2606. 14636v1 Announce Type: new Abstract: Control-function instrumental variable estimators need a first-stage residual, not merely a first-stage prediction.
By Rui Wu, Zongyuan Chen, Hong Xie, Defu Lian, Enhong Chen
arXiv:2607. 25074v1 Announce Type: cross Abstract: Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units.
By Mojtaba Eslami
arXiv:2606. 18322v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features.
By Mingyue Cui, Linghui Shen, Xingyi Yang
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.
By Vijay Erramilli
arXiv:2609.23937v1 Announce Type: cross
Abstract: Robust linear fits can resist response contamination yet remain too dense or unstable for useful global explanations. We propose penalized distillati...
By Wooyoung Shin, Seunghwan Park