arXiv Machine Learning

Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.

arXiv Machine Learning
Sep 10

Conditional Validity for Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group

The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.

By Melika Baghi
arXiv Computer Vision
Aug 28

Evaluator-Dependent Patient-Adaptive ECG Lead-Channel Allocation

The study investigates how patient‑adaptive ECG lead‑channel allocation policies perform when evaluated by different diagnostic models. Two policies, ECG‑on‑Demand and MGA, trained with a simple logistic evaluator were tested on a more powerful ResNet1D evaluator, revealing that the adaptive advantage observed with the training evaluator disappears or reverses with the stronger evaluator. Across multiple budgets, policies, and metrics, all interactions favor fixed protocols under the strong evaluator, suggesting that adaptive channel selection must be jointly optimized with the diagnostic backbone.

By Xiaoyang Li, Zeyan Tao
arXiv AI
3d ago

Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets

Signal‑Routed Temperature Scaling (SRTS‑BCE) is a 10‑parameter, argmax‑preserving calibration method that separates calibration objectives from adaptive capacity. It cross‑fits a correctness‑risk score over six logit statistics and assigns a top‑label BCE temperature to each of three risk groups, generalizing TvA‑TS when K=1. Experiments on fine‑tuned CIFAR‑100 and ViT‑B/16 show that SRTS‑BCE reduces ECE from 1.65 to 0.96 with a small calibration budget, outperforming higher‑capacity SMART+BCE when only 250 examples are available, and revealing a budget‑dependent ranking reversal on Swin‑T.

By Wenhao Liang, Liangwei Nathan Zheng, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv Machine Learning
Sep 14

Certified AI Triage of ICU Alarms

The paper introduces a certified AI triage system for ICU alarms, reframing alarm reduction as a three-way decision (retain, suppress, or defer). It demonstrates that, with a 5% budget, the system can suppress 74.8% of false ventricular‑tachycardia alarms while only silencing 1.5% of genuine ones, achieving an AUROC of 0.953 and a Challenge Score of 83.33—comparable to the best existing methods. The study also explores how grid granularity and calibration affect certification, showing that finer grids can certify fewer alarms but with tighter guarantees.

By Mohammed Sameer Syed, Rozhin Yasaei
arXiv Computer Vision
Sep 18

Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.

By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong
arXiv AI
Sep 11

Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

The paper introduces Cros, a risk‑constrained stopping layer for sequential clinical diagnosis agents that determines when to stop testing and make a diagnosis. Cros combines state‑wise error ranking, policy design on disjoint development splits, and exact tests of selective diagnostic error to provide finite‑sample guarantees. On a MIMIC‑derived abdominal‑pain benchmark, Cros achieves higher state‑error AUROC and lower selective error rates compared to baseline stopping methods, though its performance varies across development resplits.

By Yuexin Wu, Vasile Rus