arXiv AI

Reachable Global Optimization in AI Systems: How Global Is Global?

The paper "Reachable Global Optimization in AI Systems: How Global Is Global?" argues that claims of AI systems optimizing various components (prompts, policies, architectures, etc.) are underspecified unless they define the region actually reachable by the system. It introduces Reachability-Induced Optimization (RIO), a framework where a generator, verifier, controller, memory, tools, and budget determine a reachable candidate region, and proves several theoretical results about reachable-optimality and related concepts. Extensive benchmarks (66,150 trials across 270 landscapes) demonstrate that control can alter reachability, and that optimization quality, reachability quality, and control reliability must be reported separately.

arXiv AI
Aug 18

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

arXiv:2608. 15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making.

By Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv Machine Learning
1d ago

Learning Goal-Reaching Quasimetric Geometry From Finite-Time Reachability

The paper introduces ReQRL, a method for learning quasimetric geometry in goal-conditioned reinforcement learning by constraining the critic’s value gradients with finite-horizon reachability. It decouples dynamical reachability from boundary geometry, estimating both from data using state-constrained optimal control principles. Experiments on OGBench show that ReQRL matches or surpasses existing quasimetric and offline GCRL approaches.

By Daisuke Yamada, Travis Pence, Vikas Singh
arXiv AI
Aug 25

Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric

The paper introduces the Unified Path Planner (UPP), a graph‑search algorithm that balances safety and optimality by adaptively weighting heuristics and using a local inverse‑distance safety field. UPP auto‑tunes its parameters during search, guaranteeing suboptimality bounds while improving obstacle clearance. Evaluation on ten simulated environments shows UPP achieving a 0.94 OptiSafe score—significantly higher than existing methods—while adding only 0.5–1% to path length and maintaining a 100% success rate, with hardware validation on a TurtleBot confirming practical benefits.

By Jatin Kumar Arora, Soutrik Bandyopadhyay, Sunil Sulania, Shubhendu Bhasin