arXiv:2609.34879v2 Announce Type: replace
Abstract: Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Too...
By Xiang Xia, Cheng Yan, Wuyang Zhang, Fan Xu, Zhijun Fan, Shuyuan Zhang, Yanyong Zhang
RouteRepair is a method that diagnoses specific weaknesses in large language model (LLM)-generated routing heuristics by evaluating performance at the instance level and then applies targeted modifications to the heuristic components that are failing, while preserving components that already perform well. It combines routing evidence, solver behavior, and program context to set bounded repair objectives and validates each change through matched parent-child evaluation of failure recovery and collateral degradation. Experiments on the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP) show significant reductions in optimality gaps and route costs, demonstrating that failure-aware, evidence-constrained refinement can improve routing heuristics on difficult instances while maintaining performance on easier cases.
By Binghao Ji, Di Huang, Jiahui Fang, Zhiyuan Liu
arXiv:2608. 06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization.
By Lekang Jiang, Bohan Tang, Stephan Goetz, Yiwen Guo
arXiv:2608.30924v1 Announce Type: new
Abstract: Travel itinerary generation requires balancing strict spatio-temporal constraints with human preferences. Existing LLM-based planners mainly rely on st...
By Priyanshu Karmakar, Borru Vijay Sai, Shubhojit Mallick, Abhik Jana, Shreya Ghosh, Manish Gupta
arXiv:2609.36138v1 Announce Type: new
Abstract: Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or...
By Jiayi Li, Ruizhe Li
arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.
By Jiayu Liu, Cheng Qian, Zhaochen Su, Qing Zong, Shijue Huang, Bingxiang He, Yi R. Fung