arXiv AI

Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models

arXiv AI
Sep 2

Context-Grounding Gains Are Mediated by Pre-existing Machinery: Auditing GRPO, SFT, and DPO

The study investigates whether post‑training methods—GRPO, SFT, and DPO—improve language models’ ability to follow prompt evidence that conflicts with memorized knowledge. By comparing nine training variants across different scales and families, the authors find that grounding gains are modest for GRPO, moderate for Conflict‑SFT, and near‑ceiling for DPO, but all largely rely on the same causal attention‑head set present in the starting checkpoint. Removing the starting‑model grounding direction suppresses these gains, while adding it back recovers a significant portion of DPO’s improvement, indicating that existing model machinery drives most of the observed gains.

By Prakhar Gupta, Vaibhav Gupta
arXiv Machine Learning
5d ago

Decodable In-Context State and Model Output Across Training

The paper investigates how a probe can decode in‑context bindings on model errors and how probe‑guided steering can repair them. It tracks probe accuracy, model output, and steering response across public pretraining and post‑training checkpoints, noting that probe accuracy improves during Pythia pretraining and that steering benefits grow with model size. The study also shows that decoders trained on final state or candidate logits do not outperform each other on late‑checkpoint errors, and presents an information‑theoretic counterexample explaining why decodability on errors alone cannot prove discarded output information.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha
arXiv Machine Learning
Sep 11

Legible Failures: Detecting and Repairing In-Context Binding Errors

The paper investigates in-context binding errors in language models, showing that a linear probe can recover correct entity bindings from frozen hidden states even when the model outputs incorrect bindings. Across 16 checkpoints, the probe’s accuracy on failure cases surpasses a baseline by about 0.196, and a probe‑based score improves failure detection over the model’s confidence by 0.079 AUROC. Steering the residual stream toward the probe‑decoded binding further boosts accuracy by an average of 0.168 across eight models.

By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari