arXiv:2606. 12923v1 Announce Type: cross Abstract: AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects.
By Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Tim Elson
NeuronSifter is a framework for planning interventions in central nervous system microenvironments by converting treatment regimens into state‑conditional target‑occupancy fields and propagating them through microenvironment dynamics. It selects measurements based on their expected reduction in intervention loss, integrating typed outcomes into a unified posterior. In synthetic Alzheimer’s disease scenarios, occupancy conditioning improves trajectory probability scores and intervention ordering accuracy, and decision‑directed acquisition reduces terminal risk compared to a Bayesian experimental design planner.
By Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
arXiv:2606. 19831v1 Announce Type: cross Abstract: Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output.
By Hongliang Liu
arXiv:2608. 19338v1 Announce Type: cross Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions.
By Vijay Erramilli
arXiv:2607. 27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes.
By Christian Rosenthal
arXiv:2608. 01548v2 Announce Type: replace Abstract: Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--environment complex can execute, and which possibilities can be realized with finite resources.
By Yi Liu