arXiv AI By Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Tim Elson

Order Is Not Control

Read the original on arXiv AI →

arXiv:2606. 12923v1 Announce Type: cross Abstract: AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

NeuronSifter: Intervention Planning in CNS Microenvironments

NeuronSifter is a framework for planning interventions in central nervous system microenvironments by converting treatment regimens into state‑conditional target‑occupancy fields and propagating them through microenvironment dynamics. It selects measurements based on their expected reduction in intervention loss, integrating typed outcomes into a unified posterior. In synthetic Alzheimer’s disease scenarios, occupancy conditioning improves trajectory probability scores and intervention ordering accuracy, and decision‑directed acquisition reduces terminal risk compared to a Bayesian experimental design planner.

By Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
arXiv AI
Sep 4

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.

By Vijay Erramilli