arXiv:2609.34879v2 Announce Type: replace
Abstract: Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Too...
By Xiang Xia, Cheng Yan, Wuyang Zhang, Fan Xu, Zhijun Fan, Shuyuan Zhang, Yanyong Zhang
RouteRepair is a method that diagnoses specific weaknesses in large language model (LLM)-generated routing heuristics by evaluating performance at the instance level and then applies targeted modifications to the heuristic components that are failing, while preserving components that already perform well. It combines routing evidence, solver behavior, and program context to set bounded repair objectives and validates each change through matched parent-child evaluation of failure recovery and collateral degradation. Experiments on the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP) show significant reductions in optimality gaps and route costs, demonstrating that failure-aware, evidence-constrained refinement can improve routing heuristics on difficult instances while maintaining performance on easier cases.
By Binghao Ji, Di Huang, Jiahui Fang, Zhiyuan Liu
arXiv:2608. 06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization.
By Lekang Jiang, Bohan Tang, Stephan Goetz, Yiwen Guo
arXiv:2608.30924v1 Announce Type: new
Abstract: Travel itinerary generation requires balancing strict spatio-temporal constraints with human preferences. Existing LLM-based planners mainly rely on st...
By Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick, Abhik Jana, Shreya Ghosh, Manish Gupta
arXiv:2609.36138v1 Announce Type: new
Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or...
By Jiayi Li, Ruizhe Li
arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.
By Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung
arXiv:2609.13566v1 Announce Type: new
Abstract: Industrial maintenance systems involve multiple interacting assets and shared resources, making it challenging to balance reliability and operational c...
By Xian Yeow Lee, Chandrasekar Venkatraman, Ahmed Farahat
arXiv:2607. 29055v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks.
By Hanxiao Lu, Tianyi Zhang
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
arXiv:2609.05736v2 Announce Type: new
Abstract: LLM tool agents can be improved without retraining by modifying the runtime harness around a fixed model: prompts, tool interfaces, middleware, state h...
By Cen Mia Zhao, Haibo Ruan, Wenjie Chen, Pei-fen Tu, Usman Abbasi, Joel Hesch
arXiv:2509. 21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience.
By Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
The paper introduces complete cyclic subtask graphs for large language model agents, enabling a workflow controller where all subtasks are fully connected and a unified agent selects transitions based on natural‑language criteria. It evaluates task‑specific and benchmark‑generic cyclic graphs on TextCraft, ALFWorld, and Finance‑Agent, comparing them to ReAct and dependency‑directed workflows, and identifies three distinct workflow signatures that influence the effectiveness of cyclic routing. The study also provides a workflow‑signature matrix, robustness analysis, token‑cost accounting, and failure‑mode structure, concluding that cyclic subtask graphs serve as a diagnostic tool to determine when flexible backtracking is worthwhile versus when simpler controllers suffice.
By Luay Gharzeddine, Samer Saab Jr