AllocBench: Measuring Online Tool Allocation Capability in LLM Agents
arXiv:2607. 23332v2 Announce Type: replace Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse.
arXiv:2608. 13675v1 Announce Type: cross Abstract: Between October 2018 and July 2026 AI models progressed from simple systems like BERT to massive agents that solve complex math and write software.
arXiv:2607. 23332v2 Announce Type: replace Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse.
Writing Answer Set Programming (ASP) theories from scratch is a difficult and time-consuming task. We take a neurosymbolic approach to study whether a model can distill complete and correct theories, given a fixed agent harness with the solver in the loop.
arXiv:2607. 23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse.
arXiv:2605. 03195v2 Announce Type: replace Abstract: Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow responsibilities like search, debugging or terminal execution.
arXiv:2609.27717v1 Announce Type: new Abstract: Human-written agent skills encode rich workflows for real-world problem solving, but are typically used as external inference-time instructions rather...
arXiv:2608. 16620v1 Announce Type: cross Abstract: Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks.
arXiv:2606. 10933v1 Announce Type: new Abstract: LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories.
arXiv:2607. 08964v1 Announce Type: new Abstract: AI agents have become capable of autonomously completing short, well-specified tasks.
arXiv:2608. 00355v1 Announce Type: cross Abstract: Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score.
arXiv:2609.40303v1 Announce Type: new Abstract: Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnati...
arXiv:2609. 20519v1 Announce Type: new Abstract: As coding agents move from supervised code completion to unattended, around-the-clock exploration, their work expands from isolated predictions into long trajectories of reasoning, tool use, and feedback.
arXiv:2607. 14989v1 Announce Type: cross Abstract: Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tasks through interaction.