Efficient LLM Collaboration via Planning
arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achi...
arXiv:2512. 11213v2 Announce Type: replace Abstract: Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training.
arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achi...
arXiv:2511. 17006v2 Announce Type: replace Abstract: Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thinking in tokens but also acting via tool calls that directly constrain environmental interaction.
arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.
PeakBench is a new benchmark designed to evaluate how large language model agents invoke multiple tools while respecting resource constraints and parallel execution. It provides executable multi‑tool workflows with dependency annotations and measured resource profiles, and introduces a two‑part evaluation framework that separates logical planning from physical scheduling. The study shows that strong logical planning alone does not guarantee safe or efficient execution, and that providing resource information can reduce overflows and improve utilization.
arXiv:2511. 02200v2 Announce Type: replace Abstract: The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling diverse agents to integrate unique expertise, collaborate flexibly, and address challenges unattainable for individual models.
arXiv:2606. 05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions.
arXiv:2606. 15684v1 Announce Type: new Abstract: We present TickingCollabBench, a Minecraft-based multi-agent benchmark for a novel class of time-sensitive complementary collaboration tasks.
arXiv:2505. 11765v5 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications.
arXiv:2602.18998v2 Announce Type: replace Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavio...
arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.
arXiv:2607. 25816v1 Announce Type: new Abstract: Large language model agents often spend substantial wall-clock time waiting for tool call results.
GRASP is a multi-stage planning framework that separates planning into specialized modules: GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation. This strategy-aware approach yields state‑of‑the‑art accuracy on diverse datasets, outperforming direct LLM planners by up to 30.8% on ZebraLogic and reducing multi‑task degradation. GRASP’s context isolation and macro‑regularization also give it a 14.5% edge over frontier reasoning models like GPT‑5‑mini.