arXiv AI By Raj Jaiswal, Dhruv Jain, Rishabh Dhawan, Sree Krishna Uppalapati, Shin'ichi Satoh, Tanuja Ganu, Rajiv Ratn Shah

Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models

Read the original on arXiv AI →

arXiv:2607. 05199v1 Announce Type: new Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

Decoupled Physical Modeling and Execution for Physics Reasoning

The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.

By Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li, Qingyun Wang
arXiv Computation and Language
Sep 17

Code Consistency Preference Optimization Verification for Language Model Alignment

The paper introduces Code Consistency Preference Optimization Verification (CCPO), a method that generates computationally sound solutions with dependency graphs to improve execution-consistent preference optimization for language models. By building a scientific reasoning dataset and extracting reasoning steps, prerequisites, and derivability relationships, the authors compute execution consistency scores that are used to fine‑tune models such as Llama‑3‑8B and DeepSeekMath‑7B, achieving significant performance gains on MATH (+17.0%) and GSM8K (+15.1%). The extended Scientific Feasibility Control framework further boosts accuracy on PhyX physics reasoning to 50.1%, surpassing existing models while maintaining high scientific validity and reducing law violations.

By Yunlong Tan, Mingqiao Mo, Hao Zhang
arXiv AI
Aug 28

From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning

The paper investigates the reliability of rule- and model-based verifiers used in reinforcement learning with verifiable reward (RLVR) for mathematical reasoning. It finds that rule-based verifiers often miss equivalent answers in different formats, causing false negatives that degrade RL performance as models improve. Model-based verifiers achieve higher static accuracy but become vulnerable to reward hacking during RL, misclassifying certain response patterns as correct after fine-tuning.

By Yuzhen Huang, Weihao Zeng, Xingshan Zeng, Qi Zhu, Junxian He