arXiv AI By Qinglan Luo, S M A Nahian, John Guttag, S. Mazdak Abulnaga, Katie Matton

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

Read the original on arXiv AI →

The paper proposes a test‑time method to enhance the faithfulness of large language model (LLM) explanations by removing concepts not credited in the model’s explanation before re‑querying the model. This approach targets incompleteness—omissions of influential factors—rather than unsoundness, and is model‑agnostic, requiring no changes to model weights. Experiments across two datasets and multiple model families show improved faithfulness compared to standard prompting and faithfulness‑encouraging prompts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 3

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

The paper proposes a test‑time method to enhance the faithfulness of large language model (LLM) explanations by removing concepts that are not credited in the model’s explanation before re‑querying the model. This approach directly addresses the incompleteness dimension of unfaithful explanations, unlike prior methods that mainly target unsoundness. Experiments across two datasets, multiple model families, and two faithfulness metrics show that the method improves explanation faithfulness over standard prompting and faithfulness‑encouraging prompts, and it is model‑agnostic and requires no parameter changes.