arXiv Machine Learning

Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution Shifts

arXiv:2602. 02229v2 Announce Type: replace Abstract: We study the problem of monitoring model performance in dynamic environments where labeled data are limited.

arXiv Machine Learning
Jul 10

Prediction-Powered Active Testing

arXiv:2607. 08347v1 Announce Type: cross Abstract: Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled.

By Kianoosh Ashouritaklimi, Valentin Kilian, Daolang Huang, Tom Rainforth, Fran\c{c}ois Caron
Hugging Face Trending Papers
Jul 9

Prediction-Powered Active Testing

Active testing provides a label--efficient approach to risk estimation by adaptively selecting which test points should be labelled. However, existing estimators fail to exploit the informative predictions of powerful black--box models, even though such predictions are increasingly available in settings where labels remain expensive.

arXiv Machine Learning
Jul 30

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.

By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang
arXiv Machine Learning
Sep 14

PLSP (Pre-hoc Liminal Space Profiling): OOD Prediction over Detection -- An Anticipatory Approach for Machine Learning Model Reliability

The paper introduces PLSP (Pre-hoc Liminal Space Profiling), an anticipatory framework for predicting out-of-distribution (OOD) data before inference. It proposes a dataset‑independent metric called the CREDibility Score (CREDS) and introduces credibility curves and heat maps to analyze a model’s maximum credibility and behavior across datasets. Experiments on multiple datasets show that CREDS can improve model robustness to OOD prediction.

By Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman, Deepak Dhungana