arXiv:2606. 20858v2 Announce Type: replace Abstract: The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational priorities largely unexplored.
By Alan Nadelsticher Ruvalcaba
arXiv:2608. 10323v1 Announce Type: new Abstract: Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation.
By Yuxu Ge, Yifei Cheng
arXiv:2609.00129v1 Announce Type: cross
Abstract: The performance of artificial intelligence (AI) and machine learning (ML) models degrades when the problem they were trained on drifts. This is a nea...
By J. M. Diederik Kruijssen (Allora Foundation)
arXiv:2609.14418v1 Announce Type: cross
Abstract: Dynamic multi-mode resource-constrained project scheduling requires decisions to be made under precedence constraints, limited resources, multiple ex...
By Yuan Tian, Yi Mei, Mengjie Zhang
The paper proposes a single variational principle that explains how gating mechanisms, their dynamics, and neural implementations for behavioral composition can be unified. This principle yields softmax gating, an energy‑based dynamical system with guaranteed convergence, and a recurrent neural network model with context‑dependent, local interactions. Experiments across collective behavior, human decision‑making, and layered control show that the mechanism reproduces known behavioral patterns, offers interpretable accounts of behavior combination, and matches or outperforms existing methods.
By Francesca Rossi, Veronica Centorrino, Francesco Bullo, Giovanni Russo
arXiv:2607. 21971v1 Announce Type: new Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains.
By Shujin Wu, Cheng Qian, Xiusi Chen, Heng Ji
arXiv:2606. 29082v1 Announce Type: cross Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture?
By Young-Jun Lee, Seungone Kim, Minki Kang, Alistair Cheong Liang Chuen, Zerui Chen, Seungho Han, Taehee Jung, Dongyeop Kang
arXiv:2606. 09902v1 Announce Type: cross Abstract: Reservoir computing exploits the fixed dynamics of a recurrent network for temporal processing, requiring only a trained linear readout.
By Anmol Guragain, Savvas Kakalis, Juan Ignacio Godino-Llorente
arXiv:2606. 23587v2 Announce Type: replace Abstract: Previous work has found a gap between the scale of neural networks that reliably learn Conway's Game of Life, and minimal networks capable of representing the classic cellular automaton with hard-coded parameter values.
By Tashin Ahmed, Q. Tyrell Davis
arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.
By Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang
arXiv:2608. 11506v1 Announce Type: cross Abstract: Adaptive behavior under partial observability depends on internal organization that carries information beyond the current observation.
By Frederick Hayes III
arXiv:2608. 07645v1 Announce Type: new Abstract: Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks.
By Changzhi Liu, Yilun Liu, Sikuan Yan, Volker Tresp, Yunpu Ma