Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
The paper introduces topological necessities—mechanism‑invariant subgoals derived from the topology of successful trajectories—used to guide long‑horizon goal‑conditioned reinforcement learning. By computing homology in dimensions 0 and 1 over a transport‑weighted carrier, the authors obtain an enumerable gate set that forms a recursive topological gate hierarchy. These certified gates transfer across different embodiments (e.g., from PointMaze to Ant and Humanoid) without retraining, achieving state‑of‑the‑art performance on several benchmark tasks.
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
arXiv:2608. 05085v1 Announce Type: cross Abstract: Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop.
arXiv:2607. 12547v1 Announce Type: cross Abstract: We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control.
The paper introduces Network Feasibility Geometry Reinforcement Learning (NFG‑RL), a method that enforces multi‑layer network constraints—such as interference, power‑rate coupling, flow conservation, service chains, capacity, latency, and reliability—by transporting a proto‑policy through a differentiable feasibility map. By compiling heterogeneous constraints into typed residual blocks and using a variational transport operator, NFG‑RL ensures almost‑sure feasible execution and shapes exploration and gradients to respect active constraints. Experiments on two wireless‑edge surrogate environments show that NFG‑RL boosts feasible utility by 37.5–41.5 %, cuts raw‑action violations by 48.5–60.8 %, and reduces P99 delay by 57.0–75.5 % compared to leading baselines.
arXiv:2606. 00427v1 Announce Type: new Abstract: State abstraction in reinforcement learning is usually formulated as a partition of states based on reward and transition similarity.
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
arXiv:2304.10041v2 Announce Type: replace Abstract: This work investigates formal policy synthesis for continuous-state stochastic dynamic systems subject to high-level specifications expressed in li...
arXiv:2607. 15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit.
arXiv:2609.16056v1 Announce Type: cross Abstract: Humans carry behaviour knowledge of how to act in familiar situations into every new task rather than relearning it from scratch. There is no reason...
arXiv:2608.29061v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-hori...
The paper introduces Collective Counterfactual Planning (CCP), a formal model describing how teams coordinate tasks that no single member can handle alone, constrained not by capability but by representational geometry. CCP defines four critical gates—exogenous implementation coalitions, conception, consent, and task-relative verification—that determine whether a team can achieve and legitimately recognize a conjunctive goal. The authors present the Collective Counterfactual Solvability (CCS) problem, separating geometric feasibility, executable attainment, and validated completion, and provide a sound and complete four-step solvability scheme under exact representation of relay closure.
arXiv:2609.36238v1 Announce Type: new Abstract: A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Y...