arXiv AI By Mohamed Sayed, Wolfram Burgard, Tanja Katharina Kaiser

Shape Formation for the Cooperative Transportation of Arbitrary Objects Using Multi-Agent Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 09610v1 Announce Type: cross Abstract: Cooperative object transportation is essential in numerous domains, including industrial to domestic services.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 24

Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections

arXiv:2607. 21488v1 Announce Type: cross Abstract: Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) systems, which typically struggle with combinatorial action spaces, reliance on privileged information, or rigid agent designs.

By Gil Lifshits, Igal Bilik, Gilad Katz
arXiv AI
Sep 17

Learning Multi-Humanoid Pickup and Transport via Decentralized Object-Centric Control

The paper presents a decentralized, object‑centric control strategy for cooperative multi‑humanoid pickup and transport of objects with diverse sizes, weights, and shapes. Each humanoid is assigned a local attachment region on the shared object and learns to perform gripperless bimanual pinching, enabling pickup, transport, and handover without task‑specific redesign. Experiments in simulation and on real hardware demonstrate that single‑robot trained policies transfer to multi‑robot settings and that additional multi‑robot training further improves coordination.

By Bikram Pandit, Mohitvishnu S. Gadde, Aayam Kumar Shrestha, Alan Fern
arXiv Machine Learning
Sep 11

ObstaDiff: Generalizable Diffusion Policy Learning via Obstacle-aware Representations

ObstaDiff is a diffusion-policy framework that introduces a lightweight obstacle-aware visual encoder to generate structured representations of targets, obstacles, and background. By aligning these representations, the policy produces end-effector trajectories that focus on a target-centered bottleneck pose while accounting for surrounding obstacles. In real-robot greenhouse trials, ObstaDiff achieved a 75.41% task success rate and an 8.20% obstacle collision rate, outperforming existing imitation-learning baselines in cluttered agricultural settings.

By Jiawen Wang, Kevin Yao, Khalid Jawed