arXiv Machine Learning

Targeted Tests for LLM Reasoning: An Audit-Constrained Protocol

arXiv:2605. 11599v3 Announce Type: replace Abstract: Fixed reasoning benchmarks evaluate canonical prompts, but semantically valid changes in presentation can still change model behavior.

arXiv AI
Aug 18

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.

By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv AI
Aug 24

Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

The paper investigates how small lexical changes in prompts can cause large performance swings in large language models. Using a dataset of 132,000 prompt variants, the authors uncover a scaling law linking higher average task performance to lower variance and greater robustness. They identify domain-specific terminology and explicit action directives as key linguistic factors that stabilize prompts, and propose an automated Prompt-Refining Agent that reduces performance variance by 40.7% in code generation while maintaining or improving mean performance.

By Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
arXiv AI
Aug 5

Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility

arXiv:2608. 04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute.

By Mohsen Hariri, Weicong Chen, Nahal Shahini, Vikash Singh, Kai Ye, Amirhossein Samandar, Debargha Ganguly, Sreehari Sankar, Yanyan Zhang, Shouren Wang, Jerry Peng, Biyao Zhang, Michael Hinczewski, Vipin Chaudhary
arXiv Machine Learning
Sep 22

EAVer: Long-Form Factuality Verification as an End-to-End Agentic Policy

arXiv:2609.22223v1 Announce Type: cross Abstract: Long-form factuality verification is commonly implemented as a static decompose-search-verify pipeline, with separately prompted modules processing c...

By Kening Zheng, Aoying Zheng, Zhigang Chang, Yazhi Guo, Miaotian Guo, Qingwei Zong, Xianhai Xie, Weiqiang Jin, Chengze Li, Hanrong Zhang, Jie Yang, Wei-Chieh Huang, Lingzhe Zhang, Liancheng Fang, Xin Zou, Hanqian Li, Jiahao Huo, Yibo Yan, Zizhuang Deng, Lei Miao, Wei Guo, Haihong Tang, Bo Zheng, Philip S. Yu