Can AI Understand the Language of Origami?
arXiv:2603.13856v3 Announce Type: replace Abstract: Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the...
arXiv:2603.13856v3 Announce Type: replace Abstract: Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must reason about the...
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
KnowDemo is a framework that generates diverse robot demonstrations from human videos by leveraging structured manipulation knowledge. It uses a vision‑language model to extract task requirements and permissible execution variations, then resolves these against target‑scene entities to guide candidate generation and screening before motion planning. The resulting demonstrations feature multimodal behavior, alternative contact strategies, and valid subtask orders, and have been shown to improve planning success and enable sim‑to‑real policy transfer across three tasks.
arXiv:2606. 28385v1 Announce Type: cross Abstract: Recent advances in robot world models enable synthetic video generation for embodied prediction and planning.
arXiv:2606. 20118v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution.
The paper proposes a neuro‑symbolic framework that augments vision‑language‑action (VLA) models with explicit task graphs and multimodal procedural memory to handle long‑horizon manipulation tasks. Task graphs encode action dependencies, valid transitions, and branch conditions, while memory tracks the active step, completed actions, textual context, and relevant visual evidence. Human demonstrations provide spatial and temporal guidance via gaze or saliency cues, which are annotated in robot‑view teleoperation videos and used to fine‑tune VLA models. The approach is evaluated on workspace clearing and surgical‑instrument handling tasks, measuring object and destination selection, subtask completion, task progress, step‑order consistency, overall success, and procedural or execution mistakes.
arXiv:2605. 31286v2 Announce Type: replace-cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments.
AntiGrounding is a visual action-selection framework that turns short robot trajectories into both executable motion plans and rendered prompts for vision‑language model evaluation. After filtering for feasibility, each trajectory is scored on safety, task alignment, efficiency, and physical plausibility using structured multi‑view visual question answering, and the best trajectories are refined and validated by a digital twin before real‑world execution. In eight real‑world manipulation tasks, the system achieved a 71.25% success rate with a single GPT‑6 Astra evaluator, outperforming baseline methods.
arXiv:2603. 22435v2 Announce Type: replace-cross Abstract: "Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored.
RAPID is a system that automatically generates, verifies, and refines robot programs from a single visual human demonstration. It infers a testable task specification, action primitives, and an interactive environment, using an object-centric relational program representation to enable reuse beyond the demonstration. The approach was evaluated in simulation on eight contact-rich manipulation tasks and successfully deployed on a real Franka arm, showing strong generalization across object pose, shape, material, and environment.
FuncBridge is a two‑stage framework that addresses functional generalization in robotics by decoupling functional reasoning from action execution. It learns to predict generalizable 2D keypoint trajectories from action‑free data and then grounds these trajectories into robot actions with limited demonstrations. Across a benchmark of ten tools and three functions—hitting, sweeping, and hooking—FuncBridge outperforms state‑of‑the‑art methods on unseen tools in both simulation and real‑world tests.
arXiv:2608. 14047v1 Announce Type: cross Abstract: This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART).