The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.
By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu
arXiv:2609.37172v1 Announce Type: new
Abstract: Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuri...
By Jianghan Zhu, Cong Zhang, Rongjie Zhu, Chi Zhang, Zhiguang Cao
arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.
By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
arXiv:2606. 05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
By Zhenfeng Cao
The paper introduces the Universe of Universes (UoU) framework, treating the ecosystem of major large language models as a structured retrieval corpus and proposing a compositional Automated Reasoning and Machine Learning architecture for cross-model retrieval‑augmented generation. It formally defines the Benefit Yield Function (BYF), measuring marginal performance gain per added model, and identifies an implosion threshold θ* where BYF becomes zero and ensemble performance degrades. The work highlights gaps in current LLM ensemble research, such as lack of performance analysis across full model universes, and connects these findings to implications for DoD AI acquisition policy and testing of AI‑enabled systems.
By Danielle Franklin, Vasu Raj Jain
arXiv:2609.22592v1 Announce Type: cross
Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a ve...
By Aarati Andrea Noronha, Kavya Ravikumar, Carly Xiaoyu Lin