arXiv AI By Wesley Shu

Reachable Global Optimization in AI Systems: How Global Is Global?

Read the original on arXiv AI →

The paper "Reachable Global Optimization in AI Systems: How Global Is Global?" argues that claims of AI systems optimizing various components (prompts, policies, architectures, etc.) are underspecified unless they define the region actually reachable by the system. It introduces Reachability-Induced Optimization (RIO), a framework where a generator, verifier, controller, memory, tools, and budget determine a reachable candidate region, and proves several theoretical results about reachable-optimality and related concepts. Extensive benchmarks (66,150 trials across 270 landscapes) demonstrate that control can alter reachability, and that optimization quality, reachability quality, and control reliability must be reported separately.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 18

LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset

arXiv:2608. 15175v1 Announce Type: cross Abstract: Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditions and mission requirements require intelligent adaptive decision-making.

By Yousef Emami, Mohammadhossein Homaei, Hao Zhou, Miguel Guti\'errez Gait\'an, Atefeh Hajijamali Arani, Rui Zhang
arXiv Machine Learning
Aug 19

Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.

By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee