arXiv Machine Learning By Kevin Qiu, Marek Cygan

Debate2Create: Robot Co-design via Multi-Agent LLM Debate

Read the original on arXiv Machine Learning →

arXiv:2510. 25850v3 Announce Type: replace-cross Abstract: We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 12

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.

By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
arXiv AI
Sep 25

When Search Becomes Memory: Accelerating Robot Design Discovery with Self-Evolving Skills

The paper introduces Auto‑Robotist, a self‑evolving large language model (LLM) agent that transforms evolutionary robot design search traces into an explicit natural‑language skill library. Each skill records a structural archetype, evidence‑grounded rules, and supporting designs, enabling the agent to retrieve and condition LLM edits during search while still using a genetic algorithm for exploration. Experiments on seven EvoGym tasks show that Auto‑Robotist outperforms standard genetic algorithms, especially when transferring learned skills to larger design spaces.

By Yunfei Wang, Xiaohao Xu, Yang Li, Xiaonan Huang
arXiv AI
Sep 21

SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations

SynthDemo‑RL introduces a teacher‑student framework that uses an automated teacher to generate successful manipulation trajectories from simulator‑privileged state, which are then distilled into a Vision‑Language‑Action (VLA) student via supervised fine‑tuning. The student is further refined with PPO using binary task‑success rewards. On the LIBERO‑PRO benchmark, SynthDemo‑RL rescues all 27 previously unsolvable tasks and achieves near‑perfect success rates, matching performance that would otherwise require human demonstrations.

By Hiroaki Kingetsu, Hiroaki Kurihara, Kaoru Yokoo, Kenji Fukumizu, Manohar Kaul
arXiv AI
Sep 4

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

CORAL is a framework that enables autonomous multi‑agent evolution for open‑ended discovery, replacing rigid heuristics with long‑running agents that explore, reflect, and collaborate via shared memory and asynchronous execution. It incorporates safeguards such as isolated workspaces, evaluator separation, and resource management. In experiments across mathematical, algorithmic, and systems optimization tasks, CORAL achieves 3–10 times higher improvement rates with fewer evaluations than traditional evolutionary baselines, and improves the best known score on Anthropic’s kernel engineering task from 1363 to 1103 cycles.

By Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu, Xuan Jiang, Sirui Li, Cathy Wu, Bryan Kian Hsiang Low, Jinhua Zhao, Paul Pu Liang