arXiv Machine Learning

CELEUS: Certifiable and Efficient LLM Evaluation via E-Processes

arXiv:2606. 20820v2 Announce Type: replace Abstract: Can we trust evaluation scores to capture an LLM's true real-world performance?

Hugging Face Trending Papers
Aug 10

Consilience for Verifier-Free Test-Time Scaling

Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications.

arXiv Machine Learning
Aug 4

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.

By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao
arXiv Machine Learning
Aug 11

Consilience for Verifier-Free Test-Time Scaling

arXiv:2608. 09898v1 Announce Type: cross Abstract: Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts.

By Lecheng Kong, Like Hui, Haitao Mao, Jun Huan