arXiv AI By Gabriel Hurtado

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map

Read the original on arXiv AI →

arXiv:2607. 01854v1 Announce Type: cross Abstract: Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Hard-Gate Candidacy in a Deployed Validator Suite

The paper evaluates hard‑gate candidacy for validators in a deployed generative‑agent system by measuring how well each validator’s firing separates successful from failed builds. Across 13 validators and thousands of builds, only a few checks show statistically significant separation, while many fail to distinguish or never fire. The study highlights that skipped checks are recorded as passes, limiting detectable failure rates and underscoring the need for clearer evaluation records.

By Xin Xu
arXiv AI
3d ago

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage

RAGScope is a leakage‑controlled, cost‑aware protocol that evaluates evidence‑gating mechanisms for retrieval‑augmented generation (RAG) systems using only the task input, retrieved context, and answer text. The enhanced gate, RAGScope‑E, achieves an AUROC of 0.798 and an average precision of 0.660 on three RAGTruth tasks, outperforming ROUGE‑L by 0.034 in pooled AP and delivering 0.748 precision within a top‑10% review budget. It operates quickly (6.22 ms per example on CPU) and demonstrates that cheap evidence gates can effectively triage RAG outputs, though calibration must be validated and adapted for each target domain.

By Zeming Liu, Qibai Chen, Jingtao Zhang, Hang Lyu