arXiv Machine Learning

Debate2Create: Robot Co-design via Multi-Agent LLM Debate

arXiv:2510. 25850v3 Announce Type: replace-cross Abstract: We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation.

arXiv AI
Jun 12

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.

By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
arXiv AI
Sep 25

When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills

The paper introduces Auto‑Robotist, a self‑evolving large language model (LLM) agent that transforms evolutionary robot design search traces into an explicit natural‑language skill library. Each skill records a structural archetype, evidence‑grounded rules, and supporting designs, enabling the agent to retrieve and condition LLM edits during search while still using a genetic algorithm for exploration. Experiments on seven EvoGym tasks show that Auto‑Robotist outperforms standard genetic algorithms, especially when transferring learned skills to larger design spaces.

By Yunfei Wang, Xiaohao Xu, Yang Li, Xiaonan Huang
arXiv AI
Sep 21

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

SynthDemo‑RL introduces a teacher‑student framework that uses an automated teacher to generate successful manipulation trajectories from simulator‑privileged state, which are then distilled into a Vision‑Language‑Action (VLA) student via supervised fine‑tuning. The student is further refined with PPO using binary task‑success rewards. On the LIBERO‑PRO benchmark, SynthDemo‑RL rescues all 27 previously unsolvable tasks and achieves near‑perfect success rates, matching performance that would otherwise require human demonstrations.

By Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo, Kenji Fukumizu, Manohar Kaul
arXiv AI
Sep 4

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

CORAL is a framework that enables autonomous multi‑agent evolution for open‑ended discovery, replacing rigid heuristics with long‑running agents that explore, reflect, and collaborate via shared memory and asynchronous execution. It incorporates safeguards such as isolated workspaces, evaluator separation, and resource management. In experiments across mathematical, algorithmic, and systems optimization tasks, CORAL achieves 3–10 times higher improvement rates with fewer evaluations than traditional evolutionary baselines, and improves the best known score on Anthropic’s kernel engineering task from 1363 to 1103 cycles.

By Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu, Xuan Jiang, Sirui Li, Cathy Wu, Bryan Kian Hsiang Low, Jinhua Zhao, Paul Pu Liang
arXiv AI
4d ago

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

arXiv:2609.36588v1 Announce Type: cross Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLA...

By Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li
arXiv AI
Sep 7

A Schema Bounded Language Model for Refining Robot Policies Without Destabilizing Local Learning

The paper presents a decentralized navigation framework for composite heterogeneous robots that integrates a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller. Each robot independently generates and refines policies at the round level using LLM inference, while the Double DQN handles tick-level action selection based on navigation variables and LLM priors. Across 30 rounds, the full configuration achieved all goals with the lowest median completion time (42 ticks) and a 25–39% improvement over other setups.

By Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif, Carlo R. daCunha
arXiv AI
2d ago

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

arXiv:2610.02089v1 Announce Type: cross Abstract: As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use re...

By Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, Idan Szpektor, Avi Caciularu, Jongmin Park, Youngjae Yu