arXiv AI

Designing Service Systems from Textual Evidence

arXiv:2603. 10400v2 Announce Type: replace-cross Abstract: Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality control procedure.

arXiv AI
Aug 17

ASSERT: A Measurement Pipeline for GenAI Audits

arXiv:2608. 13840v1 Announce Type: cross Abstract: Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy.

By Riccardo Fogliato, Abhinav Palia, Xiawei Wang, Emily Sheng, Chad Atalla, Jean Garcia-Gathright, Nicholas Pangakis, Sharman Tan, Dan Vann, Hannah Washington, P. Alex Dow, Heba Elfardy, Hanna Wallach, Sandeep Atluri
arXiv AI
Aug 24

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

The paper introduces a framework for evaluating AI systems that not only checks final labels but also tracks the reasoning behind them through three core sources—grounds, norms, and authority—forming an eight-cell counterfactual judgment cube. It defines minimal source replacement sets, called judgment receipts, to explain changes in verdicts and provides certification cost bounds for black-box evaluators. The authors present ReasonBench, a benchmark with 19,520 cases, and demonstrate that while high standard accuracy can mask robustness issues, receipt accuracy reveals significant gaps in reasoning consistency across different models.

By Ye Chen, Weining Zhang
arXiv AI
Sep 25

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.

By Dae Woong (David), Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
arXiv AI
Sep 18

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

The paper introduces prediction‑powered smoothing (PP‑S) and its taxonomy‑aware extension (PP‑TS) to improve point and interval estimates of domain‑specific AI performance when only a limited sample of labeled units is available. It also proposes a new design‑based cross‑validation score that is approximately unbiased for selecting between direct and smoothed estimators. Experiments on a curated benchmark and real‑world agent traffic show that the proposed methods outperform direct estimators in both accuracy and coverage, and that the new score matches the performance of an independent validation sample while providing more precise error estimates.

By Sho Kawano, Zehang Richard Li, Paul A. Parker
arXiv Computation and Language
Sep 7

Auditing Bias and Safety in Voice AI Customer Care

The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.

By Vignesh Ethiraj, Ashwath David