Robots Influencing Humans to Reveal their Goals during Collaboration and Competition
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.
arXiv:2510. 17059v2 Announce Type: replace Abstract: Zero-shot imitation learning requires an agent to reproduce expert behavior from a single demonstration without additional environment interaction or gradient updates at test time.
arXiv:2502. 18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models.
The paper introduces Planning Diffusion Policy Optimization (PDPO), an offline‑to‑online reinforcement‑learning framework that employs a diffusion policy to produce short‑horizon action chunks for robot crowd navigation. PDPO is pretrained on collision‑avoidance demonstrations and fine‑tuned online with PPO, generating five‑step action sequences applied in a receding‑horizon manner. The authors also identify a benchmark artifact where agents can leave the valid domain without explicit boundary constraints, and they mitigate this by treating boundary violations as collisions, leading to improved success rates over strong baselines.
The paper introduces a method for generating legible plans in arbitrary PDDL domains by extending prior legibility research to classical planning without custom planners. It incorporates a second‑order theory of mind to estimate the observer’s perspective, enabling robots to implicitly communicate goals in human‑robot teaming. Benchmark results show that increasing legibility typically trades off with plan efficiency, and a regularizing factor is needed to balance the two.
arXiv:2509. 10656v2 Announce Type: replace-cross Abstract: For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning.