When Does a Second Model Help? Cross-Model Review in LLM Verification
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The study investigates how varying the capability of a reviewer model in a large‑language‑model (LLM) execute‑review‑revise pipeline affects rejection decisions on 100 olympiad mathematics problems. A mid‑tier reviewer improves final accuracy by 12 percentage points (from 52 % to 64 %) without damaging answers, while a self‑reviewer detects errors best (85 % recall) but rejects too often and harms correct solutions. Below a certain capability threshold the reviewer becomes inert, changing none of the answers and doubling token cost.
The paper investigates the reliability of ranking tables produced by small-sample evaluations of large language models (LLMs). Using LLM‑inferred prompt structure across eight model variants, the authors find that prompt‑structure recovery is highly unstable, with only the bottom of the ranking consistently reproducible. They demonstrate that standard evaluation practices can misrepresent model performance and propose reporting practices to improve transparency.
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
arXiv:2606. 15887v1 Announce Type: cross Abstract: Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns.
arXiv:2606. 09046v1 Announce Type: new Abstract: Useful audits reveal not only how often a model fails, but also where its failures concentrate.
arXiv:2608. 08514v1 Announce Type: new Abstract: We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across domains and models (RPC across four new task domains with Qwen3-8B, LCF across four 7-8B models).