arXiv AI By Muhammad Aziz Ullah, Abdul Serwadda

Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries

Read the original on arXiv AI →

The paper evaluates using large language model (LLM) juries to review code generated from natural language queries, focusing on MySQL text-to-SQL tasks. It benchmarks 15 open models, selects the top six, and constructs unanimous committees of varying sizes to accept a query only when all members agree. The study finds that single-model judges are inconsistent, while small unanimous committees of strong models can reduce false accepts without discarding many correct queries, and that committee composition significantly influences performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 14

Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces

arXiv:2608. 12585v1 Announce Type: new Abstract: Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong training signals during reinforcement learning, and an in-depth understanding of reasoning behaviors during model performance evaluation.

By Congchao Wang, Diwakar Singh, Qiaozi Gao, Spyros Matsoukas, Yang Liu, Mahdi Namazifar
arXiv Computation and Language
Sep 21

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

JudgeSense is a benchmark comprising 880 items from human‑labelled corpora, each presented under two differently worded instructions that ask the same question. The study evaluates 25 judges from six providers across four tasks, measuring how rewording affects agreement with the judge’s own verdicts. Results show that rewording reduces agreement on all tasks, with significant effects on two, and that stability varies across tasks and is not predicted by parameter count.

By Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang