arXiv Computation and Language By Ziang Ni, Peng Zou

Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

Read the original on arXiv Computation and Language →

The paper investigates whether large language model (LLM) agents that yield to a unanimous majority in multi‑agent debate truly change their underlying premise or merely adjust their statement. Using two‑hop factual questions with an unstated intermediate entity, the authors show that agents that concede still encode the original bridge in their internal representations, as revealed by a Jacobian‑lens analysis, even when the majority answer is wrong. Experiments across several open‑weight models demonstrate that hiding or removing the agent’s earlier answer increases conformity and that the original premise can be recovered from the question alone.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.