arXiv AI By Zeming Liu, Qibai Chen, Jingtao Zhang, Hang Lyu

RAGScope: A Leakage-Controlled, Cost-Aware Evidence-Gating Protocol for RAG Hallucination Triage

Read the original on arXiv AI →

RAGScope is a leakage‑controlled, cost‑aware protocol that evaluates evidence‑gating mechanisms for retrieval‑augmented generation (RAG) systems using only the task input, retrieved context, and answer text. The enhanced gate, RAGScope‑E, achieves an AUROC of 0.798 and an average precision of 0.660 on three RAGTruth tasks, outperforming ROUGE‑L by 0.034 in pooled AP and delivering 0.748 precision within a top‑10% review budget. It operates quickly (6.22 ms per example on CPU) and demonstrates that cheap evidence gates can effectively triage RAG outputs, though calibration must be validated and adapted for each target domain.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

AutoTuneBench introduces a trustworthy measurement protocol for evaluating how large language model agents auto‑tune GPU kernels and serving engines. The benchmark addresses four failure modes—strawman baselines, machine‑dependent timing, saturated tasks, and infrastructure defects—by enforcing code‑frozen protocols, database validation, anti‑cheat checks, pre‑registered comparisons, and external result anchoring. Using this protocol, the authors demonstrate that previously reported speedups are inflated, revealing more modest improvements across different engines and machines.

By Li Chen