Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
The paper introduces HALTER, a graph-based system that automates the reset and evaluation of long-horizon robot manipulation tasks. HALTER constructs a spatial scene graph from point clouds and vision models, uses an LLM to score rollouts, plan resets, and verify success, all without labeled success images. In experiments on a Franka arm, HALTER restores scenes in 76% of episodes, improves skill completion estimation, and reduces operator time by 72% compared to manual reset.
arXiv:2608. 16889v1 Announce Type: cross Abstract: Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task.
The paper introduces URAI, a Universal Robot‑Agent Interface that separates robot control into two roles: a programming agent that writes reusable, task‑specific tools from intent, and an execution agent that calls these tools in a feedback loop. This design keeps high‑level decision making in the model while delegating low‑level motion to code, allowing tool revisions to persist across episodes without retraining the foundation model. Experiments on RoboDojo and AgileX tasks show significant gains in success rate, speed, and token efficiency compared to direct fingertip control and pre‑written programs.
arXiv:2609.39304v1 Announce Type: cross Abstract: When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posin...
The paper investigates task collapse—a failure mode where online RL fine‑tuning of a pretrained flow‑matching vision‑language‑action policy erodes performance on individual tasks—using a 450M‑parameter SmolVLA policy on LIBERO‑10. Three exploration‑noise strategies are compared: a fixed noise scale, a learned noise network, and an uncertainty‑gated controller that reallocates exploration based on novelty and competence signals without task labels. The uncertainty‑gated controller prevents task collapse across all tested seeds, whereas the other two approaches consistently cause collapse, demonstrating its effectiveness in preserving task performance during fine‑tuning.
The paper investigates whether closed‑loop robot software generated and refined by a coding agent can be reused to acquire policies for new tasks. For each source task, the agent creates policy code from a few demonstrations, iteratively improves it with simulation feedback, and stores the validated implementations. When applied to new tasks, the agent uses these archived implementations, additional demonstrations, and execution feedback to produce a final policy that runs without further model calls, achieving higher success rates than starting from scratch or from unoptimized source code.
The paper introduces neuro‑symbolic computer use, a method that learns reusable policies to execute recurring computer workflows efficiently. Instead of re‑planning each run, the learned policy encodes stable decisions (ordering, variables, loops, branches) into executable code while delegating observation‑dependent decisions to neural models. Using neuro‑symbolic policy iteration, the approach iteratively refines the policy from a single agent trajectory, diagnoses failures, and revises the code with a coding model, achieving superior Pass^3 scores and significant reductions in per‑run cost and latency on OSWorld‑Verified and ScienceBoard benchmarks.
arXiv:2608. 12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state.
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly master the individual skills, yet the chain still fails: err...
arXiv:2603.01209v3 Announce Type: replace Abstract: In CodeAct, language-model agents write Python that calls tools and use execution feedback to choose actions. Persistent runtimes preserve Python v...
arXiv:2607. 17136v1 Announce Type: cross Abstract: Agentic computer-use RL is reported in single runs, and those numbers mislead.