arXiv Machine Learning

Global Sequential Testing for Multi-Stream Auditing

arXiv:2602. 21479v3 Announce Type: replace-cross Abstract: Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed.

arXiv Machine Learning
Jun 5

Multi-Armed Sequential Hypothesis Testing by Betting

arXiv:2603. 17925v2 Announce Type: replace-cross Abstract: We consider a variant of sequential testing by betting where, at each time step, the statistician is presented with multiple data sources (arms) and obtains data by choosing one of the arms.

By Ricardo J. Sandoval, Ian Waudby-Smith, Michael I. Jordan
arXiv Machine Learning
Jul 7

Knowing When to Stop: Predicting Execution-Consistency Convergence in Text-to-SQL

arXiv:2607. 03991v1 Announce Type: new Abstract: Repeated LLM calls are the standard way to estimate how trustworthy a Text-to-SQL result is: run the pipeline multiple times, judge each SQL execution, and use the consistency of the verdicts as a confidence signal.

By Yaron Anavi, Mor Aisenberg, Nadav Nesher, Elena Khabibullina, Isabella Cattinelli