arXiv:2511.08097v2 Announce Type: replace-cross
Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...
By Dheeraj Narasimha, Nicolas Gast
arXiv:2609.07247v1 Announce Type: new
Abstract: Scaling the interaction horizon-the maximum number of environment interactions per episode-improves LLM agents on long-horizon tasks, and curriculum-ba...
By Gangyi Zhang, Junjie Meng, Letian Zhang, Wei Wu, Yang Zheng, Dong Wang, Yang Liu, Guanjun Jiang, Chongming Gao
arXiv:2607. 26370v1 Announce Type: cross Abstract: We propose a self-adaptive online learning for control method for tracking unknown target dynamics.
By Atharva Navsalkar, Hongyu Zhou, Vasileios Tzoumas
arXiv:2606. 09929v1 Announce Type: cross Abstract: Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable.
By Caleb Munigety
arXiv:2604. 26256v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-tailed trajectories indispensable for model performance block the entire training pipeline.
By Tianhao Hu, Xiangcheng Liu, Yuchun Miao, Youshao Xiao, Hongyu Zang, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Yu Yang, Yi-Kai Zhang, Yueqing Sun, Chengcheng Han, Xiandi Ma, Wei Wang, Qi Gu, Yerui Sun, Yuchen Xie, Xunliang Cai
The paper introduces a conflict‑predictive variable horizon for distributed model predictive control in multi‑drone collision avoidance. Each drone locally estimates future conflicts using observed positions and confidence funnels, then selects the minimal horizon that covers the farthest predicted conflict, shrinking in clear air and expanding only when necessary. The authors prove that this adaptive horizon preserves recursive feasibility and asymptotic stability for linear models, and demonstrate in simulation that it reduces per‑step solver cost and total computation while maintaining separation on dense benchmarks.
By Linda M\"{u}mken, Michael Schwung, Stefan Lier, Andreas Schwung