arXiv AI

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

arXiv:2607. 18960v1 Announce Type: cross Abstract: Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth acquiring.

Hugging Face Trending Papers
Jul 21

SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement

Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth acquiring. We present \sys{}, a statistics-first gating architecture that treats procurement as a cost-aware routing problem over three intrinsic quality axes -- diversity, utility, and redundancy.

arXiv Computation and Language
Sep 25

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains. whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."

By Delip Rao, Chris Callison-Burch
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
Hugging Face Trending Papers
Sep 24

JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places

The study compares Jev, a typed classifier that outputs probabilities over allowed answers, with three flash-tier LLM rubric judges across nine panels from seven benchmarks. Jev’s accuracy differs significantly from an LLM judge in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons are inconclusive. In terms of cost and speed, Jev is 29 to 325 times cheaper and 30 to 220 times faster than the LLM judges, and a cascade that defers uncertain Jev verdicts to an LLM yields only modest gains of up to 2.0 points over the best single judge.

arXiv AI
Aug 24

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

JuryProbe is an empirical diagnostic tool designed to assess consensus risk in panels of reference‑free large language model judges used for factuality verification. It estimates risk by measuring false‑negative correlations and false‑consensus lift from a labeled calibration probe, and routes high‑risk majority decisions to judges with trusted references. The approach was validated on FEVER corruptions, showing that flagged decisions can be grounded without additional reference acquisition in most cases, while reducing false accepts by about 0.4% and avoiding 28% of reference acquisitions.

By Tianxin Zhou, Ruixi Lin