arXiv Machine Learning

NeuronSifter: Intervention Planning in CNS Microenvironments

NeuronSifter is a framework for planning interventions in central nervous system microenvironments by converting treatment regimens into state‑conditional target‑occupancy fields and propagating them through microenvironment dynamics. It selects measurements based on their expected reduction in intervention loss, integrating typed outcomes into a unified posterior. In synthetic Alzheimer’s disease scenarios, occupancy conditioning improves trajectory probability scores and intervention ordering accuracy, and decision‑directed acquisition reduces terminal risk compared to a Bayesian experimental design planner.

arXiv Machine Learning
Aug 11

Learning Multi-Timescale Interventions under Safety and Resource Constraints

arXiv:2508. 03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them.

By David Mguni, Wanrong Yang, Jing Dong, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
arXiv AI
Jun 12

Order Is Not Control

arXiv:2606. 12923v1 Announce Type: cross Abstract: AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects.

By Gareth Seneque, Lap-Hang Ho, Nafise Erfanian Saeedi, Jeffrey Molendijk, Tim Elson
arXiv Machine Learning
Sep 10

Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.

By Melika Baghi
arXiv Machine Learning
Sep 18

CARE-VI: Conservative Adaptive Reliability Estimation for Value Improvement in Off-Policy Actor-Critic Learning

CARE‑VI introduces a framework for improving value targets in off‑policy actor‑critic learning by combining Conservative Adaptive Ranking and Screening (CARS), Selector‑Evaluator Value Assessment (SEVA), and Dynamic Adaptive Risk‑aware Enhancement (DARE). CARS limits candidate actions to a budgeted prefix and expands it only when uncertainty exceeds a threshold; SEVA orders candidates with selector critics and reviews their values with an evaluator critic, capping the value at the selector reference; DARE adjusts residual corrections based on candidate reliability and signal gaps. Theoretical analysis bounds errors in each component, and empirical tests on SAC, TD3, and TD7 across four MuJoCo tasks show CARE‑VI consistently outperforms baselines in mean return.

By Xiang Zou, Shengzhu Shi, Junqi Gao, Zhichang Guo
arXiv AI
Aug 26

Revelation Control

Revelation Control studies how to price interventions that reveal hidden state only when the revealed distinctions can alter a consequential decision, while separately accounting for any useful progress the intervention itself creates. The authors develop a framework for learning systems that defines decision‑sufficient revelation, revelation depth, and a cost‑adjusted factorization criterion, and they provide a target‑independent protocol for model‑specific instantiation. Experiments on Qwen2.5‑7B and Mistral‑7B‑v0.3 show that deeper future‑learning probes have positive decision value and that productive reuse yields strict equal‑compute utility advantages, supporting a structural transfer of the decision theory and evaluation protocol across architectures.

By Qinyou Wang
arXiv Computer Vision
Sep 18

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.

By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong