arXiv Machine Learning By Rui Wu, Zongyuan Chen, Hong Xie, Defu Lian, Enhong Chen

Projective Graph Residualization: Variation-Allocation Frontiers for Control-Function IV

Read the original on arXiv Machine Learning →

arXiv:2606. 14636v2 Announce Type: replace Abstract: Control-function instrumental-variable estimators pass an estimated first-stage residual to an outcome model.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.

By Vijay Erramilli