arXiv:2602. 16793v2 Announce Type: replace Abstract: In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO).
By Xingyu Dang, Rohit Agarwal, Rodrigo Porto, Anirudh Goyal, Liam H Fowl, Sanjeev Arora
Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof quality along five...
arXiv:2606. 07810v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability.
By Anish Laddha, Nitesh Pradhan, Gaurav Srivastava
ProofJudge is an agentic large language model that evaluates the quality of formal proofs in Lean 4 beyond mere correctness. It scores proofs on five dimensions—library leverage, automation fit, structural clarity, statement quality, and Mathlib conventions—using tool access to the relevant repository commit. The system was tested on 218 Mathlib declarations, achieving alignment with human reviewers between 63.5% and 80.8% and releasing its harness, dataset, and traces for open research.
By Shane Caldwell
arXiv:2607. 01223v1 Announce Type: new Abstract: When should an AI system's answer be trusted?
By Ben Slivinski, Michael Saldivar
Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears creative is truly original and effective is hard to establis...
arXiv:2608. 15979v1 Announce Type: new Abstract: Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules.
By Eric Xie, Wenqian Ye, Aidong Zhang
The study investigates whether Jev, a typed classifier that outputs probabilities over allowed answers without generating text, can replace large language model (LLM) rubric judges. Across nine panels from seven benchmarks, Jev’s accuracy differed significantly from LLM judges in only 8 of 27 paired comparisons, performing best on binary criteria and worse only on graded ones, while most other comparisons were inconclusive. In terms of cost and speed, Jev was 29 to 325 times cheaper and 30 to 220 times faster than the flash‑tier LLM judges, and a cascade approach that defers uncertain Jev verdicts to an LLM yielded only modest gains.
whyItMatters":"The findings suggest that a lightweight classifier like Jev can serve as an efficient first‑stage evaluator, potentially reducing the reliance on expensive and slow LLM judges in automated grading pipelines."
By Delip Rao, Chris Callison-Burch
arXiv:2609.39229v1 Announce Type: cross
Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
By Elia Onofri, Roberto Di Pietro
arXiv:2602. 20629v3 Announce Type: replace Abstract: As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation.
By Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye, Edward Zhang, Ruslans Aleksejevs, Todor Anti\'c, Polina Baron, Sujeet Bhalerao, Shubhrajit Bhattacharya, Zachary Burton, John Byrne, Hyungjun Choi, Nujhat Ahmed Disha, Koppany Istv\'an Encz, Yuchen Fang, Robert Joseph George, Ebrahim Ghorbani, Alan Goldfarb, Jing Guo, Meghal Gupta, Stefano Huber, Annika Kanckos, Minjung Kang, Hyun Jong Kim, Dino Lorenzini, Levi Lorenzo, Tianyi Mao, Giovanni Marzenta, Ariane M. Masuda, Lukas Mauth, Ana Mickovic, Andres Miniguano-Trujillo, Antoine Moulin, Wenqi Ni, Tomos Parry, Kevin Ren, Hossein Roodbarani, Mathieu Rundstr\"om, Manjil Saikia, Detchat Samart, Rebecca Steiner, Connor Stewart, Dhara Thakkar, Jeffrey Tse, Vasiliki Velona, Yunhai Xiang, Sibel Yal\c{c}{\i}n, Jun Yan, Ji Zeng, Arman Cohan, Quanquan C. Liu
arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.
By Junbo Jacob Lian, Huiling Chen, Hanzhang Qin, Chung-Piaw Teo
arXiv:2609.26550v3 Announce Type: replace
Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge tha...
By Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman