arXiv AI By Yuyang Dai, Xueqing Peng, Lingfei Qian, Zhuohan Xie

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

Read the original on arXiv AI →

arXiv:2606. 17459v1 Announce Type: new Abstract: Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.