arXiv:2607. 04572v1 Announce Type: new Abstract: Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does not guarantee that the answer was derived from the student-facing problem.
By Bonan Shen, Dingyan Shang, Youting Wang, Tao Ning
The paper proposes a test‑time method to enhance the faithfulness of large language model (LLM) explanations by removing concepts not credited in the model’s explanation before re‑querying the model. This approach targets incompleteness—omissions of influential factors—rather than unsoundness, and is model‑agnostic, requiring no changes to model weights. Experiments across two datasets and multiple model families show improved faithfulness compared to standard prompting and faithfulness‑encouraging prompts.
By Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton
arXiv:2606. 24267v1 Announce Type: cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing.
By Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques
arXiv:2606. 24267v2 Announce Type: replace-cross Abstract: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing.
By Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques
The paper investigates whether giving AI monitors access to the final answer improves their ability to verify reasoning. Using 237 step‑by‑step solutions to physics exam questions, the authors found that answer access mainly helps monitors detect inconsistencies with the final answer rather than independently checking the reasoning. Certification of the answer increased overall accuracy and error localization but reduced the ability to flag critical traces where the answer was correct but the reasoning was flawed.
By Will Yeadon, Sergio Ju\'arez, Paul Mackay, T. J. Dowling, Elise Agra, Oto-obong Inyang, Arin Mizouri, Craig P. Testrow
arXiv:2607. 22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time.
By Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati