The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.
By Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li, Qingyun Wang
arXiv:2609.23367v1 Announce Type: cross
Abstract: FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising...
By Bakar Chargeishvili
The paper introduces Code Consistency Preference Optimization Verification (CCPO), a method that generates computationally sound solutions with dependency graphs to improve execution-consistent preference optimization for language models. By building a scientific reasoning dataset and extracting reasoning steps, prerequisites, and derivability relationships, the authors compute execution consistency scores that are used to fine‑tune models such as Llama‑3‑8B and DeepSeekMath‑7B, achieving significant performance gains on MATH (+17.0%) and GSM8K (+15.1%). The extended Scientific Feasibility Control framework further boosts accuracy on PhyX physics reasoning to 50.1%, surpassing existing models while maintaining high scientific validity and reducing law violations.
By Yunlong Tan, Mingqiao Mo, Hao Zhang
The paper investigates the reliability of rule- and model-based verifiers used in reinforcement learning with verifiable reward (RLVR) for mathematical reasoning. It finds that rule-based verifiers often miss equivalent answers in different formats, causing false negatives that degrade RL performance as models improve. Model-based verifiers achieve higher static accuracy but become vulnerable to reward hacking during RL, misclassifying certain response patterns as correct after fine-tuning.
By Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He
arXiv:2605. 02909v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs).
By Kazuki Egashira, Mark Vero, Jasper Dekoninck, Florian E. Dorner, Robin Staab, Martin Vechev
arXiv:2608. 11573v1 Announce Type: cross Abstract: Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs).
By Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan
arXiv:2606. 07108v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking".
By Tengyao Tu, Yulin Li, Hui-Ling Zhen, Libo Qin, Zhoujun Wei, Jinghua Piao, Zhuotao Tian, Yong Li, Min Zhang
arXiv:2602. 10576v2 Announce Type: replace-cross Abstract: Symbolic regression aims to distill mathematical equations from observational data.
By Boxiao Wang, Kai Li, Tianyi Liu, Chen Li, Junzhe Wang, Yifan Zhang, Jian Cheng
The paper introduces a verifier‑guided explainable reasoning framework for educational question answering that integrates gold‑anchored QLoRA, a task‑aware symbolic router, and group‑relative RLVR. It adapts Qwen2.5‑3B‑Instruct with field‑weighted QLoRA supervision, routes logic problems to a FOL/Z3 verifier and physics problems to a symbolic solver, and uses verifier feedback for candidate evaluation, self‑revision, and reward construction. Experiments on 438 held‑out examples show that RLVR boosts reasoning depth (P3) from 50.68 % to 72.20 %, while symbolic verification improves answer reliability at the system level.
By Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo, Minh Khang Tran, Duy Phuong Tran
arXiv:2604. 23333v2 Announce Type: replace Abstract: Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability.
By Liaoyaqi Wang, Chunsheng Zuo, William Jurayj, Benjamin Van Durme, Anqi Liu
arXiv:2606. 04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as biology, chemistry, and physics remains largely unexplored.
By Xiangyu Zhao, Hengyuan Zhao, Yiheng Wang, Wanghan Xu, Yuhao Zhou, Qinglong Cao, Zhiwang Zhou, Lei Bai, Wenlong Zhang, Xiao-Ming Wu
PhysElite is a new bilingual multimodal benchmark designed to evaluate large language models on Olympiad-level physics problems. It contains 11,586 problems, each paired with visual diagrams, step-by-step bilingual Chinese‑English solution derivations, and the final answer. Benchmarking 18 models revealed that even the best reaches only 33.7% accuracy, and a step‑level analysis highlights where models falter in reasoning.
By Ruoran Xu, Wending Gao, Liyunfeng Chen, Aixin Shi, Haoyu Cheng, Zixiang Fang, Yiqiang Zou, Qiufeng Wang