arXiv AI
The paper "Reachable Global Optimization in AI Systems: How Global Is Global?" argues that claims of AI systems optimizing various components (prompts, policies, architectures, etc.) are underspecified unless they define the region actually reachable by the system. It introduces Reachability-Induced Optimization (RIO), a framework where a generator, verifier, controller, memory, tools, and budget determine a reachable candidate region, and proves several theoretical results about reachable-optimality and related concepts. Extensive benchmarks (66,150 trials across 270 landscapes) demonstrate that control can alter reachability, and that optimization quality, reachability quality, and control reliability must be reported separately.
arXiv:2608.29397v1 Announce Type: new
Abstract: Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is...
By Zixiang Xu, Jiaan Wang, Fandong Meng
arXiv:2608. 15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making.
By Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
arXiv:2607. 00064v1 Announce Type: new Abstract: As technology advances, many path-planning algorithms have been proposed for Air Traffic Management, yet their operational adoption in tactical control remains limited, revealing a misalignment between algorithmic design priorities and air traffic controllers' needs.
By Yiyuan Zou, Wenying Lyu, Clark Borst
arXiv:2605. 30719v2 Announce Type: replace-cross Abstract: We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.
By Stephane Hatgis-Kessell, Emma Brunskill
arXiv:2503. 01236v3 Announce Type: replace-cross Abstract: This paper addresses fixed-graph terrain-aware path refinement, in which a global planner is restricted to a predefined route space and may remain optimal within that space while missing lower-cost terrain corridors available in the native-resolution map.
By Ling Xiao, Toshihiko Yamasaki
Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.
By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv:2508. 19186v2 Announce Type: replace-cross Abstract: Reactive obstacle avoidance methods often cause agents to become trapped in local minima, because they can often only reason one step ahead (i.
By Christopher Chandler, Bernd Porr, Giulia Lafratta, Alice Miller
The paper introduces ReQRL, a method for learning quasimetric geometry in goal-conditioned reinforcement learning by constraining the critic’s value gradients with finite-horizon reachability. It decouples dynamical reachability from boundary geometry, estimating both from data using state-constrained optimal control principles. Experiments on OGBench show that ReQRL matches or surpasses existing quasimetric and offline GCRL approaches.
By Daisuke Yamada, Travis Pence, Vikas Singh
arXiv:2606. 02016v1 Announce Type: new Abstract: Algorithm Selection (AS) aims to automatically identify the most suitable optimization algorithm for a given problem instance by leveraging measurable problem characteristics and historical performance data.
By Gjorgjina Cenikj, Jakub Kudela, Eva Tuba, Tome Eftimov
The paper introduces the Unified Path Planner (UPP), a graph‑search algorithm that balances safety and optimality by adaptively weighting heuristics and using a local inverse‑distance safety field. UPP auto‑tunes its parameters during search, guaranteeing suboptimality bounds while improving obstacle clearance. Evaluation on ten simulated environments shows UPP achieving a 0.94 OptiSafe score—significantly higher than existing methods—while adding only 0.5–1% to path length and maintaining a 100% success rate, with hardware validation on a TurtleBot confirming practical benefits.
By Jatin Kumar Arora, Soutrik Bandyopadhyay, Sunil Sulania, Shubhendu Bhasin
arXiv:2608. 06702v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets under strict real-time constraints.
By Vaibhav Sanjay, Jiaoyang Li
Mobile robots that operate in side by side with humans and critical facilities must reach their goals at low cost, despite often unknown true traversal costs of the map apriori and imperfect actuation...