arXiv:2604.14251v2 Announce Type: replace
Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be es...
By Edoardo Pona, Milad Kazemi, Mehran Hosseini, Yali Du, David Watson, Osvaldo Simeone, Nicola Paoletti
arXiv:2608. 14089v1 Announce Type: new Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves.
By Thiago Sandoval, Ufuk Topcu
arXiv:2608. 04045v1 Announce Type: cross Abstract: Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data.
By Chinmoy Mitra, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Rakibul Islam, M. F. Mridha
arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.
By Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
arXiv:2503. 15581v2 Announce Type: replace Abstract: Real-time safety assessment is critical for ensuring the reliable operation of complex dynamic systems.
By Songqiao Hu, Zeyi Liu, Lufeng Hao, Yinzhong Cheng, Xiao He
arXiv:2607. 19361v1 Announce Type: cross Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm.
By Sanjay Mishra, Divya Chukkapalli, Ganesh R. Naik
The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.
By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao
arXiv:2601. 16406v2 Announce Type: replace-cross Abstract: Rare-event prediction is critical in domains such as healthcare, finance, reliability engineering, customer support, aviation safety, where positive outcomes are infrequent yet potentially catastrophic.
By Vitaly Bulgakov, Alexander Turchin
arXiv:2607. 28665v1 Announce Type: cross Abstract: Automated driving systems (ADSs) are becoming ubiquitous.
By Bidhya Shrestha, Christos Papadopoulos
Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry without sharing raw data. This study examines two complementary challenges: benign heterogeneity, where honest operators observe different operating conditions and fault modes, and adversarial heterogeneity, where compromised operators submit poisoned updates.
The paper introduces Confounding-Valid Counterfactual Conformal Inference (CV‑CCI), a method that merges abundant observational telemetry with limited randomized data to answer network operators’ ‘what‑if’ questions about key performance indicators (KPIs). CV‑CCI uses the General Synthetic‑Powered Inference principle to maintain finite‑sample coverage guarantees even when hidden confounding is present, while producing tighter prediction sets than existing baselines. Experiments on two radio access network control tasks demonstrate the method’s validity under hidden confounding and its improved efficiency.
By Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin
The paper presents a machine‑learning approach (ML_CP) that automatically learns patterns and rules for complex event prediction, reducing reliance on manual rule creation. It incorporates sensitivity analysis to assess how output varies with each input and uses conformal prediction to generate uncertainty‑aware prediction intervals. Experiments on binary, multi‑level classification, and regression tasks show promising results for safety‑critical embedded systems.
By Maria J. P. Peixoto, Akramul Azim