The paper introduces a new benchmark that evaluates large language models (LLMs) on their agentic mathematical reasoning rather than just final answers. It aligns problem‑solving behaviors with a taxonomy of reusable mathematical atomic capabilities and includes planning, action, and feedback tasks in both textual and multimodal settings. Experiments show that models with similar end‑to‑end accuracy can have very different agentic profiles, highlighting the importance of process‑level evaluation.
By Jiayi Kuang, Yinghui Li, Yunze Song, Keyu Chen, Zhifeng Shen, Yangning Li, Yidong Wang, Di Yin, Ruizhi Qiao, Xing Sun, Kai Jin, Ying Shen, Liang Lin, Philip S. Yu
arXiv:2609.37172v1 Announce Type: new
Abstract: Large language models (LLMs) have emerged as powerful tools for automated heuristic design (AHD), enabling iterative generation and refinement of heuri...
By Jianghan Zhu, Cong Zhang, Rongjie Zhu, Chi Zhang, Zhiguang Cao
arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.
By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
arXiv:2606. 05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
By Zhenfeng Cao
The paper introduces the Universe of Universes (UoU) framework, treating the ecosystem of major large language models as a structured retrieval corpus and proposing a compositional Automated Reasoning and Machine Learning architecture for cross-model retrieval‑augmented generation. It formally defines the Benefit Yield Function (BYF), measuring marginal performance gain per added model, and identifies an implosion threshold θ* where BYF becomes zero and ensemble performance degrades. The work highlights gaps in current LLM ensemble research, such as lack of performance analysis across full model universes, and connects these findings to implications for DoD AI acquisition policy and testing of AI‑enabled systems.
By Danielle Franklin, Vasu Raj Jain
arXiv:2609.22592v1 Announce Type: cross
Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a ve...
By Aarati Andrea Noronha, Kavya Ravikumar, Carly Xiaoyu Lin
arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.
By Kabir Moghe, Peter Chin
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness.
arXiv:2606. 05608v2 Announce Type: replace-cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
By Zhenfeng Cao
arXiv:2602. 11198v2 Announce Type: replace-cross Abstract: Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding assistants can generate correct, framework-specific code.
By Shafiuddin Rehan Ahmed, Sourabh Deshpande
arXiv:2510. 15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke.
By Pavan C Shekar, Aswanth Krishnan
The paper introduces GEARS, a framework that treats ranking optimization as an autonomous discovery process within a programmable experimentation environment. By encapsulating ranking expert knowledge into reusable agent skills, GEARS allows operators to steer systems through high-level intent personalization rather than static model selection. The framework also includes validation hooks to enforce statistical robustness and filter out brittle policies, and experimental results show that GEARS consistently finds near‑Pareto‑efficient policies while maintaining deployment stability.
By Longfei Yun, Yihan Wu, Haoran Liu, Xiaoxuan Liu, Ziyun Xu, Yi Wang, Yang Xia, Pengfei Wang, Mingze Gao, Yunxiang Wang, Changfan Chen, Wenjie Fu, Hong Yan, Junfeng Pan