arXiv AI By De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma

Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation

Read the original on arXiv AI →

arXiv:2608. 04408v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

TISD: On-Policy Self-Distillation with Trajectory Intervention

The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.

By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
arXiv Machine Learning
Sep 22

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

The paper presents a causal analysis of a compressed VLA policy that performs well in offline tests but fails in closed‑loop execution on a simulated pick‑and‑place task. An 8‑layer distillation of Octo‑Base retains most parameters and passes all offline metrics, yet collapses during deployment, with early stages degrading gradually and final transport failing entirely. The failure is traced to a negative, late‑heavy residual in the action trace, and standard remedies (continued training, offline data, command‑level compensation, clamping) do not restore performance; only a minimal‑pair intervention that mixes deployment‑distribution rollouts with teacher data restores parity with the teacher. whyItMatters":"The study demonstrates that offline validation metrics alone are insufficient to guarantee closed‑loop success for compressed policies, highlighting the need for targeted deployment‑time testing and interventions."

By Fengze Jia (The Ohio State University)