arXiv AI

Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

The paper presents a graph‑constrained agentic framework that enables large language models to adapt retail supply‑chain decision modules to evolving requirements. It jointly selects intervention routes and admissible module changes, validating candidates against downstream KPIs. Experiments with 100 warehouse requirements and three LLMs show the framework improves end‑to‑end success from 72–76% to 79–83%.

arXiv AI
Aug 18

Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback

arXiv:2608. 15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve.

By Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
arXiv AI
Sep 10

Learning to Configure Agentic AI Systems

The paper introduces ARC, a lightweight hierarchical policy that learns to configure LLM‑based agent systems on a per‑query basis by treating each configuration as a temporally extended option in a semi‑Markov decision process. Unlike fixed templates or hand‑tuned heuristics, ARC dynamically selects workflows, tools, token budgets, and prompts tailored to the difficulty of each query. Experiments on reasoning, tool‑use, and agentic benchmarks show that ARC outperforms budget‑matched tool‑augmented LLMs, boosting reasoning accuracy by 31.3%, tool‑use accuracy by 13.95%, and doubling success on the τ‑Bench Airline Pass task from 9.0% to 18.0%.

By Aditya Taparia, Som Sagar, Ransalu Senanayake
arXiv AI
Jul 21

Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning

arXiv:2607. 17331v1 Announce Type: new Abstract: Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human specialists, because classical rule-based automation cannot reason about exceptions and monolithic AI assistants degrade when asked to coordinate across functional boundaries.

By Zhihao Liu, Tianyu Wang, Xi Vincent Wang, Lihui Wang
arXiv AI
Aug 3

DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain

arXiv:2605. 07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations.

By Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou
arXiv AI
Jul 14

Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems

arXiv:2607. 10768v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently.

By Yongchang Fu, Xinjie Huang, Chengjun Dai, Chengzhe Feng, Junshao Zhang, Hong Zhu
arXiv AI
Jun 9

Exploring Autonomous Agentic Data Engineering for Model Specialization

arXiv:2605. 30407v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.

By Yujie Luo, Xiangyuan Ru, Jingsheng Zheng, Jingjing Wang, Yuqi Zhu, Jintian Zhang, Runnan Fang, Kewei Xu, Ye Liu, Zheng Wei, Jiang Bian, Zang Li, Shumin Deng