Assessing AI in Introductory Physics Problem Solving
arXiv:2607. 14303v1 Announce Type: cross Abstract: Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving.
arXiv:2512. 10785v3 Announce Type: replace-cross Abstract: Generative AI offers new opportunities for individualized and adaptive learning, e.
arXiv:2607. 14303v1 Announce Type: cross Abstract: Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving.
arXiv:2504.02323v5 Announce Type: replace Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored variou...
The paper describes a pilot study of a generative AI practice platform designed to give immediate, scaffolded feedback to students in a large mathematics class. By moving human oversight to the verification of solutions rather than real‑time grading, the platform aims to reduce delays while maintaining trust and accountability. The study investigates student engagement, perceived value, and reliability of the AI feedback, and explores how these findings might apply to other engineering subjects.
The paper introduces a verifier‑guided explainable reasoning framework for educational question answering that integrates gold‑anchored QLoRA, a task‑aware symbolic router, and group‑relative RLVR. It adapts Qwen2.5‑3B‑Instruct with field‑weighted QLoRA supervision, routes logic problems to a FOL/Z3 verifier and physics problems to a symbolic solver, and uses verifier feedback for candidate evaluation, self‑revision, and reward construction. Experiments on 438 held‑out examples show that RLVR boosts reasoning depth (P3) from 50.68 % to 72.20 %, while symbolic verification improves answer reliability at the system level.
arXiv:2609.22553v1 Announce Type: new Abstract: Effective LLM tutoring depends on correctly identifying the specific error in a student's reasoning before generating feedback. We study this problem i...
The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).
Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.
arXiv:2609.13152v1 Announce Type: new Abstract: Large language models (LLMs) perform strongly on static science benchmarks, yet their ability to reason about the physical world through active experim...
arXiv:2607. 05199v1 Announce Type: new Abstract: Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows.
arXiv:2608. 07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc.
arXiv:2606. 04751v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents in scientific tasks.
The paper argues that user feedback from real interactions is a valuable learning signal for Large Language Models (LLMs), contrary to recent claims that it is too noisy to use. By creating synthetic data with a clear ground truth and testing on naturalistic data, the authors show that revisions guided by user feedback fix targeted issues more often than baseline revisions. They further reveal that current evaluation methods bias against feedback‑driven improvements, as judges tend to overlook genuinely corrected responses and favor inferior baselines.