Guaranteed Adaptive Modality Acquisition: When the Policy Chooses Its Own Calibration Group
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
The paper introduces a certified AI triage system for ICU alarms, reframing alarm reduction as a three-way decision (retain, suppress, or defer). It demonstrates that, with a 5% budget, the system can suppress 74.8% of false ventricular‑tachycardia alarms while only silencing 1.5% of genuine ones, achieving an AUROC of 0.953 and a Challenge Score of 83.33—comparable to the best existing methods. The study also explores how grid granularity and calibration affect certification, showing that finer grids can certify fewer alarms but with tighter guarantees.
arXiv:2608. 15520v1 Announce Type: new Abstract: A multimodal system may begin inference holding only some of its inputs and may acquire the rest at a cost.
arXiv:2606. 15153v1 Announce Type: new Abstract: Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha.
The paper introduces RouteCert, a method for ensuring risk control in multimodal systems that acquire inputs adaptively. It shows that conditional calibration can remain valid even when the acquisition policy determines the calibration group, and provides two finite‑sample constructions: threshold‑free routing with terminal‑pattern calibration and simultaneous validation of policy‑pattern pairs. Experiments on a clinical ECG task and masked multimodal benchmarks demonstrate that RouteCert achieves low disagreement rates and competitive answered fractions while validating each acquisition stage separately.
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
arXiv:2608.27808v1 Announce Type: cross Abstract: Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. O...
arXiv:2605. 30837v2 Announce Type: replace-cross Abstract: Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable.
The paper introduces a claim‑safe protocol for evaluating closed‑loop AI systems, consisting of three actions: Refuse, Decompose, and Refresh. It demonstrates the protocol in a simulator with 24 policy components and 1,440 held‑out cases, showing that abstention and stable false admission rates are low while providing detailed statistical diagnostics. The approach emphasizes that evaluation results should be tied to observable support and statistical calibration rather than a single PASS/FAIL label.
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
The paper introduces a method to certify selective prediction in machine learning systems by computing the availability of safety gates through exact-binomial inversion and dynamic programming. It demonstrates that a truth-informed planner can significantly improve mean coverage over naive approaches, and that reallocating error budgets further enhances coverage across diverse applications such as LLM tool‑calling, content moderation, lesion classification, and recommendation. The study highlights the importance of planning and finite‑sample estimation in ensuring reliable, granular deployment of selective predictors.
The paper introduces the twin‑prefix framework to evaluate how the size of the verification unit—i.e., how many actions a pre‑execution LLM monitor reviews in one call—affects its performance. By pairing each gold plan with a twin that differs by a single write and injecting a controlled error, the authors isolate the impact of review length on catch rates and false rejections. Their findings show that longer review windows increase rejection rates but do not improve discrimination, with the highest informedness occurring at one or two actions across all judges and domains.
The study investigates how patient‑adaptive ECG lead‑channel allocation policies perform when evaluated by different diagnostic models. Two policies, ECG‑on‑Demand and MGA, trained with a simple logistic evaluator were tested on a more powerful ResNet1D evaluator, revealing that the adaptive advantage observed with the training evaluator disappears or reverses with the stronger evaluator. Across multiple budgets, policies, and metrics, all interactions favor fixed protocols under the strong evaluator, suggesting that adaptive channel selection must be jointly optimized with the diagnostic backbone.
arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.