arXiv:2609.17325v1 Announce Type: new
Abstract: Biological cells can be viewed as individual, interacting agents whose collective dynamics give rise to adaptive behaviour at multiple levels of organi...
By Anatoly Belikov
The paper introduces a self‑evolving scientific agent that uses large language models and iterative code generation to build interpretable, physically‑reasoned white‑box controllers for complex systems. The agent deploys candidate controllers in simulations, diagnoses dynamic behavior from multimodal evidence, and refines source code until a robust policy is achieved. Applied to a nonlinear fluid‑structure interaction problem—a two‑joint dogfish swimmer navigating an unsteady wake—the agent autonomously designs a controller that consistently reaches targets across a wide range of conditions without retraining.
By Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
arXiv:2607. 18719v1 Announce Type: cross Abstract: This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after learning and enables uninstructed agents to implicitly complement the overall work based on the actions of other agents.
By Yamato Takahagi, Gentoku Nakasone, Yoshinari Motokawa, Toshiharu Sugawara
arXiv:2606. 08405v1 Announce Type: new Abstract: While data-intensive deep reinforcement learning can optimize complex control policies, scientific discovery in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control architectures.
By Boai Sun, Wenjin Guo, Zongmin Yu, Liu Yang
arXiv:2604. 12645v2 Announce Type: replace-cross Abstract: Although autonomous underwater vehicles promise the capability of marine ecosystem monitoring, their deployment is fundamentally limited by the difficulty of controlling vehicles under highly uncertain and non-stationary underwater dynamics.
By Melvin Laux, Yi-Ling Liu, Rina Alo, S\"oren T\"opper, Mariela De Lucas Alvarez, Frank Kirchner, Rebecca Adam
arXiv:2508. 14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces.
By Thomas Carta, Cl\'ement Romac, Loris Gaven, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier
The paper introduces a closed‑loop framework using autotelic reinforcement learning to explore and manipulate complex systems, specifically Lenia, a continuous cellular automaton. An agent called CARL autonomously samples diverse goals and learns a goal‑conditioned policy that intervenes with minimal, local perturbations. CARL demonstrates three key abilities: discovering stable solitons more efficiently than heuristic baselines, steering existing solitons with few interventions, and enabling humans to guide solitons through maze environments in real time via high‑level commands. The agents generalize zero‑shot to out‑of‑distribution conditions, suggesting a path toward artificial experimentalist agents that can discover and control emergent phenomena.
By Marko Cvjetko, Benedikt Hartl, Michael Levin, Cl\'ement Moulin-Frier, Pierre-Yves Oudeyer
Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.
By Chang Nie, Zhe Liu, Hesheng Wang
AUSO (Action-level Unified Skill Optimization) is a method that unifies skill learning and skill use through a progressive, action-aware optimization process. It starts by jointly learning from teacher guidance and environmental outcomes, then shifts to outcome-based policy optimization, and finally evaluates each action under skill-conditioned and skill-free contexts to strengthen beneficial skill-sensitive actions while suppressing harmful ones. Experiments on ALFWorld, WebShop, and SearchQA demonstrate that AUSO consistently improves agent performance and out-of-distribution generalization compared to competitive baselines.
By Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
arXiv:2606. 12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance.
By Haoyuan Deng, Yitong Gao, Yudong Lin, Haichao Liu, Zhenyu Wu, Ziwei Wang
arXiv:2601. 17454v2 Announce Type: replace-cross Abstract: Centralized value learning underlies a broad class of multi-agent reinforcement learning methods, but its claimed advantage is typically evaluated in settings that confound coordination structure with function approximation and partial observability.
By Muhammad Ahmed Atif, Nehal Naeem Haji, Mohammad Shahid Shaikh, Muhammad Ebad Atif
CORAL is a framework that enables autonomous multi‑agent evolution for open‑ended discovery, replacing rigid heuristics with long‑running agents that explore, reflect, and collaborate via shared memory and asynchronous execution. It incorporates safeguards such as isolated workspaces, evaluator separation, and resource management. In experiments across mathematical, algorithmic, and systems optimization tasks, CORAL achieves 3–10 times higher improvement rates with fewer evaluations than traditional evolutionary baselines, and improves the best known score on Anthropic’s kernel engineering task from 1363 to 1103 cycles.
By Ao Qu, Han Zheng, Zijian Zhou, Yihao Yan, Yihong Tang, Shao Yong Ong, Fenglu Hong, Kaichen Zhou, Chonghe Jiang, Minwei Kong, Jiacheng Zhu, Xuan Jiang, Sirui Li, Cathy Wu, Bryan Kian Hsiang Low, Jinhua Zhao, Paul Pu Liang