arXiv Machine Learning

Adaptive state-action abstractions via rate-distortion

arXiv:2606. 06123v1 Announce Type: new Abstract: When learning to walk, infants seem to address a coarse version of the problem first - stay upright, reach the caregiver - and refine it only when further practice at that resolution stops paying off.

arXiv AI
2d ago

Measuring the Stability Assumption Behind Action Chunking

The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.

By Aryan Goyal
arXiv AI
Sep 25

Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning

The paper introduces MI‑SARSA, an on‑policy temporal‑difference algorithm that incorporates mutual‑information regularization to model bounded rationality in reinforcement learning. By penalizing state‑specific deviations from a learned marginal action prior, the algorithm selectively uses state information only when the expected return outweighs the informational cost, yielding a reward‑complexity tradeoff. MI‑SARSA also predicts reaction times, showing that stronger information penalties lead to simpler policies, lower control costs, and faster responses, while regularization mitigates performance loss after environmental shifts at the expense of asymptotic return.

By James Wu, Chris R. Sims
arXiv Machine Learning
Sep 10

Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions

The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.

By Defu Cao, Angela Zhou
arXiv AI
2d ago

Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

The paper introduces Generalized Implicit Temporal Abstraction (GITA), a method for goal-conditioned reinforcement learning that conditions a single value function on multiple temporal abstraction levels (k). By aggregating advantage-weighted supervision across various k values, GITA preserves both long-range signal and local resolution without committing to a single k. Experiments on OGBench show that GITA outperforms existing offline GCRL baselines, improving average success rates by 25 percentage points over HIQL and 7 percentage points over OTA.

By Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado