arXiv AI By Xingguo Chen, Zhaohui Wu, Jinguo Ye, Chao Li, Shangdong Yang, Guang Yang, Skylar Liang, Wenhao Wang

Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

Read the original on arXiv AI →

The paper introduces Regularized Emphatic Temporal‑Difference Learning (RETD), a modification of ETD that normalizes the post‑shock dynamics while preserving the emphatic TD signal and importance ratios. RETD achieves almost‑sure convergence under harmonic diminishing stepsizes and provides a conditional constant‑stepsize moment‑contraction guarantee, demonstrating negative Lyapunov exponents on a two‑state counterexample and the Baird point. Extensive experiments confirm RETD’s ability to recover the ETD fixed point, exhibit a non‑monotone stability region, and maintain task‑dependent performance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 3

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

arXiv:2606. 03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- which governs the \emph{temporal coupling} between teacher and student -- has not been systematically studied as a stability variable.

By Haowei Guo, Baolong Bi, Ruicheng Zhang, Bingqian Sun, Wentao Zhang
arXiv Machine Learning
Sep 25

Intrinsic-Extrinsic Coupling in Learning Dynamics

The paper introduces a framework for intrinsic‑extrinsic coupling in learning dynamics, defining it via a continuation‑conditioned value of a constrained learning‑state intervention and observation‑relative fibers. It presents an executable finite‑frame classifier‑head that protects current logits while repairing historical margins, and distinguishes local admissibility, intervention value, and complete‑policy performance. Experiments on CLINC‑derived class‑incremental tasks, output distillation with RoBERTa, and SGDW dynamics demonstrate that coupling can produce both positive and negative interactions, and that coordinated content controls can match or exceed development gains while guided allocation reduces cross‑entropy loss compared to standard replay.

By Qinyou Wang
arXiv Machine Learning
Sep 14

Adapt or Forget: Provable Tradeoffs Between Adam and SGD in Nonstationary Optimization

The paper presents a theoretical study of Adam in non‑stationary stochastic optimization, distinguishing two regimes: Euclidean tracking under adaptive strong monotonicity and high‑probability projected stationarity for general smooth objectives. It derives finite‑time bounds that decompose into initialization, objective drift, first‑moment tracking error (β₁), and preconditioner perturbation (β₂), and characterizes burn‑in times for constant and step‑decay schedules. The analysis reveals a noise–drift tradeoff, showing that in noise‑dominated settings Adam’s adaptive mechanisms can improve guarantees, while in drift‑dominated settings they may worsen tracking, potentially making vanilla SGD preferable.

By Sharan Sahu, Abir Sarkar, Cameron J. Hogan, Martin T. Wells