arXiv AI

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

arXiv:2606. 29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading.

arXiv AI
Sep 10

Agentic ML Exploration (A-MLE) for Ads Ranking

arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...

By Erwin Gao, Vinodh Kumar Sunkara, Jingyi Guan, Qinjin Jia, Hangjun Xu, Xiang Ji, Sherman Wong, Surya Teja Chavali, Pratik Vaishnavi, Aryan Pandhi, Xiaoyu Deng, Zhaodong Wang, Samarth Inani, Fan Yang, Jakob Moberg, Zoe Zu, Nicolas Bievre, Sami Khenissi, Amit Jaspal, Ehsan Fakharizadi, Srinidhi Viswanathan, Dorothy Sun, Abishek Vanam, Sneha Iyer, Sheela Yadawad, Wenjie Chen, Gaby Nahum, Junhua Gu, Peter Chu, Yucheng Liu, Xin Zhao, Vitor Cid, Chaorong Chen, Vijay Pappu, Ashwin Kumar, Wenlin Chen, Ben Schulte, Deepak Chandra, Ritwik Tewari
arXiv AI
Jun 17

Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation

arXiv:2606. 17459v1 Announce Type: new Abstract: Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus on isolated cognitive tasks such as reasoning, knowledge retrieval, and economic rationality in stylized settings.

By Yuyang Dai, Xueqing Peng, Lingfei Qian, Zhuohan Xie
arXiv AI
Jul 28

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications

arXiv:2607. 23124v1 Announce Type: new Abstract: Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and interaction settings.

By Hao Jiang, Gangtao Xin, Yingdi Huang, Guojie Zhu, Jiangshan Zhang, Xinyuan Lin, Yunkun Xu, Chengyu Shen, Wenlong Fei, Jiawei Li, Yujie Fu, Sichen Kang, Tingyu Xie, Yedi Hu, Jingren Zhang, Hongcheng Gao, Jianshu Zeng, Chong Chen, Chang Guo, Chao Feng, Feng Wang, Fulin Lin, Jinchao Ma, Lang Mei, Li Huang, Liyan Liu, Qing He, Shuting Tao, Siyu Mo, Xiangnan Chen, Xiaohan Yu, Xiaoyang Li, Yanheng Hou, Yanyu Wu, Zhihan Yang, Wentao Zhang, Yang Gao, Zhao Cao
arXiv AI
Jul 14

Can Agentic Trading Systems Pay for Their Own Intelligence?

arXiv:2607. 10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value.

By Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
arXiv AI
Aug 20

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities and youth, set lineups, and respond to a board that can fire it, all using 26 tools and roughly 340–400 decision stops, with a deterministic engine producing a final score without human or LLM judges. The benchmark includes a solo track where each of 15 frontier models competes against a frozen scripted world, and an Arena track where the same models plus a scripted anchor share one 20‑year world, allowing the first head‑to‑head evaluation at this scale. whyItMatters":"FM‑Bench provides a rigorous, large‑scale test of sustained, cumulative decision‑making in language‑model agents, revealing that managerial strategy—not computational scale or vendor—drives performance over long horizons."

By Tianyou Wang, Chongyang Gao, Kezhen Chen, Chen Dong, Yinghao He, Donghan Li, Wangcheng Xu, Hongjiu Zhang, Chi Li
Hugging Face Trending Papers
Aug 19

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision‑making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities, set lineups, and respond to a board that can fire it, all while a deterministic engine aggregates the outcomes into a final score without human judgment. The benchmark evaluates six behavioral capabilities and compares 15 frontier models in solo and arena tracks, revealing that managerial behavior—not computational scale—drives performance.