arXiv AI By Prakhar Gupta, Vaibhav Gupta

Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

Read the original on arXiv AI →

The study investigates whether post‑training methods—GRPO, SFT, and DPO—improve language models’ ability to follow prompt evidence that conflicts with memorized knowledge. By comparing nine training variants across different scales and families, the authors find that grounding gains are modest for GRPO, moderate for Conflict‑SFT, and near‑ceiling for DPO, but all largely rely on the same causal attention‑head set present in the starting checkpoint. Removing the starting‑model grounding direction suppresses these gains, while adding it back recovers a significant portion of DPO’s improvement, indicating that existing model machinery drives most of the observed gains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.