SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.
By Shaoliang Yang, Jun Wang
arXiv:2605.10894v2 Announce Type: replace
Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner h...
By Moritz Stammel, Fabio De Sousa Ribeiro, Raghav Mehta, M\'elanie Roschewitz, Ben Glocker
arXiv:2606. 10066v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) are evaluated on public benchmarks whose images and question-answer pairs have been freely downloadable for years, yet reported accuracy assumes these examples were absent from pretraining.
By Bruce Changlong Xu, Lan Wu, Alexander Ryu
arXiv:2607.24852v4 Announce Type: replace
Abstract: Internal self-consistency cannot certify the accuracy of a photogrammetric reconstruction, and the failure is structural rather than a matter of tu...
By Behnam Asadi
arXiv:2608.29604v1 Announce Type: cross
Abstract: Vision-language retrieval with CLIP-style dual encoders achieves strong cross-modal performance, yet practical accuracy often hinges on localized sem...
By Siyi Liu, Xiaorong Zhu, Enjun Du, Xinyu Zuo, Lisheng Duan, Haijin Liang, Jin Ma, Junfu Pu, Yongqi Zhang
The paper introduces PACT, a provenance‑conserving fusion method for typed action admission in human‑robot collaboration. PACT treats evidence countability as a relational variable, using a supplied provenance partition to define countable units and accumulating support only across these units. Experiments on 31,200 evaluations in 48 scene clusters show that PACT achieves a lower normalized risk‑coverage area than singleton aggregation, and in offline human‑robot collaboration it admits 47 of 57 reference‑consistent candidates without any reference‑inconsistent admissions.
By Zekai Jin, Hanrong Zhang, Yihong Tang, Fei Hu, Zhen Dong, Yi Shao
arXiv:2606. 15153v1 Announce Type: new Abstract: Selective prediction with distribution-free risk control promises that, with confidence 1-delta over the calibration draw, the error rate of accepted inputs stays below a user budget alpha.
By Jingwen Zhou, Mingzhe Wang
A surveillance camera is an image sensor whose silent physical degradation invalidates every downstream consumer of its data. In-situ integrity alarms for such vision sensors require low false-alarm rates, bounded computation, and diagnosable behavior under nuisance illumination changes.
arXiv:2607. 27069v2 Announce Type: cross Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts.
By Feixiang Liu, Qiang Qiu, Lanbo Sun, Nan Wei, Huawei Shen, Xueqi Cheng
arXiv:2607. 12278v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks.
By Wenhao Zhang, Zhongliang Zhou, John Kang, Sheng Li
arXiv:2606. 07865v1 Announce Type: cross Abstract: Scientific machine learning is limited less by model size than by the data it is trained on.
By Daniel N. Wilke
arXiv:2606. 21806v2 Announce Type: replace Abstract: Deep generative models reproduce the observational distribution of their training data, inheriting any spurious associations it contains.
By Jingyuan Chen, Kangrui Ruan, Junzhe Zhang