Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv:2607. 28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agentic AI, particularly in separating inference, orchestration, and execution layers for autonomous AI agents.
By Konstantinos I. Roumeliotis, Ranjan Sapkota
arXiv:2607. 15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration.
By Shuaiyu Zhou, Fengpeng Yue, Zengjie Hu, Yuanzhe Shen, Chenyang Zhang, feng hong, Cao Liu, Ke Zeng
arXiv:2607. 06233v1 Announce Type: new Abstract: LLM-powered data agents are playing an increasingly important role in data-driven decision making.
By Ziting Wang, Yin Li, Zuhao Yang, Xiuchang Li, Jiale Bai, Gao Cong
arXiv:2606. 31229v1 Announce Type: new Abstract: Ideation plays a pivotal role in scientific discovery.
By Keyu Zhao, Lingyan Kong, Fengli Xu, Yong Li
arXiv:2608. 03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet.
By Jiayu Cao, Xingyuan Zeng, feiyu Li, Zhijing Huang, Xujie Yuan, Rongxiang Chen, Shimin Di, Libin Zheng, Jian Yin
MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.
By Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Gongshen Liu, Zhuosheng Zhang
arXiv:2609.01045v1 Announce Type: new
Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities as powerful components in agentic systems, enabling sophisticated reasoning and...
By Enci Zhang, Haofeng Wang, Yuesheng Zhu, Xiaole Cui, Guibo Luo
arXiv:2606. 11070v1 Announce Type: cross Abstract: Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems.
By Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin, Swasthi P Rao, Shikhhar Siingh, Houhan Lu, Nadia Bathaee, Sriharsha Hatwar, Paresh Dashore, Anmol Jain, Kshitij Tayal, Xiuzhu Lin, Anirban Das, Sambit Sahu, Shi-Xiong Zhang
arXiv:2510. 15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke.
By Pavan C Shekar, Aswanth Krishnan
arXiv:2607. 00627v1 Announce Type: new Abstract: Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a static context - does not reliably produce persistent, manipulable representations of an external world.
By Alexey Potapov
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings.