arXiv AI

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors

arXiv:2607. 04572v2 Announce Type: replace Abstract: Large language model (LLM) tutors may have access to teacher notes, answer keys, rubrics, or retrieved solutions while producing student-facing explanations.

arXiv AI
Sep 7

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

The paper proposes a test‑time method to enhance the faithfulness of large language model (LLM) explanations by removing concepts not credited in the model’s explanation before re‑querying the model. This approach targets incompleteness—omissions of influential factors—rather than unsoundness, and is model‑agnostic, requiring no changes to model weights. Experiments across two datasets and multiple model families show improved faithfulness compared to standard prompting and faithfulness‑encouraging prompts.

By Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton
arXiv AI
Sep 2

The Answer Is Not the Argument

The paper investigates whether giving AI monitors access to the final answer improves their ability to verify reasoning. Using 237 step‑by‑step solutions to physics exam questions, the authors found that answer access mainly helps monitors detect inconsistencies with the final answer rather than independently checking the reasoning. Certification of the answer increased overall accuracy and error localization but reduced the ability to flag critical traces where the answer was correct but the reasoning was flawed.

By Will Yeadon, Sergio Ju\'arez, Paul Mackay, T. J. Dowling, Elise Agra, Oto-obong Inyang, Arin Mizouri, Craig P. Testrow
Hugging Face Trending Papers
Sep 3

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

The paper proposes a test‑time method to enhance the faithfulness of large language model (LLM) explanations by removing concepts that are not credited in the model’s explanation before re‑querying the model. This approach directly addresses the incompleteness dimension of unfaithful explanations, unlike prior methods that mainly target unsoundness. Experiments across two datasets, multiple model families, and two faithfulness metrics show that the method improves explanation faithfulness over standard prompting and faithfulness‑encouraging prompts, and it is model‑agnostic and requires no parameter changes.