Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
arXiv:2608.21748v1 Announce Type: new Abstract: Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds ar...
The study investigates how patient‑adaptive ECG lead‑channel allocation policies perform when evaluated by different diagnostic models. Two policies, ECG‑on‑Demand and MGA, trained with a simple logistic evaluator were tested on a more powerful ResNet1D evaluator, revealing that the adaptive advantage observed with the training evaluator disappears or reverses with the stronger evaluator. Across multiple budgets, policies, and metrics, all interactions favor fixed protocols under the strong evaluator, suggesting that adaptive channel selection must be jointly optimized with the diagnostic backbone.
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
arXiv:2609.20700v1 Announce Type: new Abstract: Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflat...
Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, wit...
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
arXiv:2607. 26752v1 Announce Type: new Abstract: Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning.
The paper argues that cross‑view correspondence, commonly used in agent evaluation and trace‑based learning, functions as a measurement intervention. Removing or altering this correspondence can create artificial sensitivity or invariance, and multiple optimal correspondences can obscure mechanism labels and learning credit. The authors propose a validity theory with two‑sided validation, all‑optima identification, and uncertainty propagation, and demonstrate through experiments that unvalidated correspondences can misattribute credit and erase harmful responses.
The paper introduces Cros, a risk‑constrained stopping layer for sequential clinical diagnosis agents that determines when to stop testing and make a diagnosis. Cros combines state‑wise error ranking, policy design on disjoint development splits, and exact tests of selective diagnostic error to provide finite‑sample guarantees. On a MIMIC‑derived abdominal‑pain benchmark, Cros achieves higher state‑error AUROC and lower selective error rates compared to baseline stopping methods, though its performance varies across development resplits.
arXiv:2606. 20115v3 Announce Type: replace Abstract: Conformal risk control (CRC) provides distribution-free segmentation guarantees by calibrating a prediction-set threshold on held-out data.
arXiv:2606. 20115v1 Announce Type: new Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-out data.