arXiv Machine Learning By Jan Malte T\"opperwien, Aditya Mohan, Marius Lindauer

Beyond Success Rates: Trainability and Extractability for Offline GCRL

Read the original on arXiv Machine Learning →

arXiv:2602. 05459v2 Announce Type: replace Abstract: Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 4

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

arXiv:2609. 03241v1 Announce Type: cross Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode.

By Zixun Huang, Kishan Panaganti, Haitao Mi, Leowei Liang
arXiv AI
1d ago

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.

By Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado
arXiv Machine Learning
Jun 25

Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

arXiv:2606. 26080v1 Announce Type: new Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale.

By Changdae Oh, Wendi Li, Seongheon Park, Samuel Yeh, Tanwi Mallick, Sharon Li
arXiv Machine Learning
Jun 16

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

arXiv:2606. 17043v1 Announce Type: cross Abstract: When pretrained VLA policies are fine-tuned through online RL, each rollout episode produces only a single binary outcome (success or failure), yet the actor update requires per-transition supervision.

By Tongyan Fang, Siyuan Huang, Naiyu Fang, Ganlong Zhao, Zhongjin Luo, Jianbo Liu, Xiaogang Wang, Ying Dong, Hongsheng Li