Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
The study investigates whether post‑training methods—GRPO, SFT, and DPO—improve language models’ ability to follow prompt evidence that conflicts with memorized knowledge. By comparing nine training variants across different scales and families, the authors find that grounding gains are modest for GRPO, moderate for Conflict‑SFT, and near‑ceiling for DPO, but all largely rely on the same causal attention‑head set present in the starting checkpoint. Removing the starting‑model grounding direction suppresses these gains, while adding it back recovers a significant portion of DPO’s improvement, indicating that existing model machinery drives most of the observed gains.
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
arXiv:2609.36569v1 Announce Type: cross Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...
arXiv:2605. 07284v2 Announce Type: replace Abstract: A late-layer change learned during post-training may work on the base model's earlier state, or it may depend on earlier computation learned with it.
arXiv:2609.38465v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that...
arXiv:2605. 12705v2 Announce Type: replace Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning?
arXiv:2609.06107v1 Announce Type: new Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...
arXiv:2610.00767v1 Announce Type: cross Abstract: Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later trai...
Measuring training data influence consistently across language model pretraining is challenging. It is difficult to select downstream tasks or validation sets representative of a model's general capabilities, and reliance on task performance at intermediate checkpoints complicates comparisons across training.
arXiv:2606. 04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT).
arXiv:2607. 16244v1 Announce Type: cross Abstract: Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit.
arXiv:2606. 18487v1 Announce Type: cross Abstract: The standard heuristic of selecting the SFT checkpoint with the highest pass@1 for GRPO can fail when SFT compresses the rollout distribution.
arXiv:2607. 12640v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supervised checkpoint in the hope of producing a stronger agent.