arXiv AI By Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

Read the original on arXiv AI →

arXiv:2607. 28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.