The paper investigates whether response safety can be measured by the cosine similarity between a response embedding and the mean embedding of known‑safe responses. Using four frozen encoders and prompt‑controlled datasets, the authors find that a simple prototype (mean safe embedding) performs poorly (ROC‑AUC 0.457‑0.545) while an explicit safe‑minus‑unsafe reference achieves higher scores (0.588‑0.738). The study shows that a class mean is merely a location, not a safety direction, and that a reference with sufficient unsafe mass is needed to orient safety judgments.
By Sahil Kadadekar
arXiv:2606. 29054v1 Announce Type: new Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and standard heuristic abstention policies miss user-specified risk targets by 7.
By Varun Kotte
arXiv:2608. 03172v1 Announce Type: new Abstract: Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S.
By Qiming Bao, Sherry J. H. Feng, Kim Chester Eugenio, Meng Fon
arXiv:2604. 11305v3 Announce Type: replace Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR).
By Meiyi Zhu, Osvaldo Simeone
arXiv:2607. 27712v1 Announce Type: new Abstract: Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is position-agnostic.
By Angshuman Chakravertty, Rahul Maheshwari
The paper investigates how many normal samples are required to reliably set an alarm threshold for few‑shot anomaly detectors, focusing on distribution‑free certification limits. Using a frozen DINOv2 PCA residual ranker on 15 MVTec and 12 VisA categories, the authors show that simple leave‑one‑image‑out calibration is limited by resolution and shift, leading to empirical false‑alarm rates far above the nominal level. They derive a category‑count feasibility calculus, demonstrating that at least 14, 29, and 59 independent category draws are needed for 95% upper confidence bounds at α=0.20, 0.10, and 0.05, and propose the CRESS protocol to split source categories into reference, proposal, and certification roles.
whyItMatters":"The study provides concrete numerical thresholds for the amount of source evidence needed to guarantee reliable anomaly detection in new categories, informing practical deployment of few‑shot detectors."
By Gia Huy Thai, Nguyen Thai Anh
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
arXiv:2608. 14617v1 Announce Type: cross Abstract: A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation, sequential Bayesian odds updating, Dempster-Shafer combination, and conformal prediction) into one pipeline.
By Surya Saka
The paper introduces a new taxonomy for benchmark contamination that categorizes leakage by the mitigation it defeats—direct, derivative, temporal, distributional, and acquired—covering both training‑time and evaluation‑time scenarios. It proposes a four‑field disclosure protocol to record contamination status alongside benchmark scores, and provides a JSON schema, validator, and examples. An empirical study of 41 documents using a pre‑registered instrument shows limited reporting of contamination types and variable reliability, highlighting gaps in current disclosure practices.
By Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato
arXiv:2606. 10154v1 Announce Type: new Abstract: Quantized checkpoints are often screened first with quality metrics and only later, if at all, with direct safety tests.
By Sahil Kadadekar
arXiv:2606. 20115v3 Announce Type: replace Abstract: Conformal risk control (CRC) provides distribution-free segmentation guarantees by calibrating a prediction-set threshold on held-out data.
By Nafis Fuad Shahid
arXiv:2606. 20115v1 Announce Type: new Abstract: Conformal risk control (CRC) provides distribution-free guarantees on segmentation quality by calibrating a prediction-set threshold on held-out data.
By Nafis Fuad Shahid