arXiv AI

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.

arXiv Computation and Language
Sep 24

Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents

The paper introduces WebMRE, an offline benchmark comprising 541 tasks and 5,293 steps extracted from WebArena trajectories, designed to provide deterministic scoring for web agents without live environments. It enables the first systematic study of how guide sentences and grounded actions reinforce each other, showing that jointly decoding a guide improves element selection accuracy and that the guide acts as a causal instruction channel. The authors fine‑tune models that outperform leading zero‑shot baselines on all offline metrics.

By Chengguang Gan, Yunhao Liang, QingHao Zhang, Shiwen Ni
arXiv AI
Jul 3

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

arXiv:2607. 01480v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout against a verifier and updates the policy from that episode-level signal.

By Ye Liu, Srijan Bansal, Bo Pang, Yang Li, Zeyu Leo Liu, Yifei Ming, Zixuan Ke, Shafiq Joty, Semih Yavuz
arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos