AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
arXiv:2511. 02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability.
arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.
arXiv:2512. 11213v2 Announce Type: replace Abstract: Scaling test-time computation has been shown to significantly improve large language model (LLM) performance without additional training.
arXiv:2606. 19787v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR) work remains unclear.
arXiv:2511. 17006v2 Announce Type: replace Abstract: Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thinking in tokens but also acting via tool calls that directly constrain environmental interaction.
arXiv:2608. 05519v1 Announce Type: new Abstract: Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic.
arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achi...
arXiv:2609.01600v1 Announce Type: cross Abstract: Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a lo...
GRASP is a multi-stage planning framework that improves the reliability of large language models on complex tasks. It separates planning into three specialized modules—GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation—allowing context isolation and strict macro-regularization. Experiments show GRASP outperforms direct LLM planners by significant margins on datasets such as Natural Plan Calendar Scheduling, ZebraLogic, and SciBench Math, and it mitigates performance collapse in multi-task and dual-task settings.
arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.
GRASP is a multi-stage planning framework that separates planning into specialized modules: GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation. This strategy-aware approach yields state‑of‑the‑art accuracy on diverse datasets, outperforming direct LLM planners by up to 30.8% on ZebraLogic and reducing multi‑task degradation. GRASP’s context isolation and macro‑regularization also give it a 14.5% edge over frontier reasoning models like GPT‑5‑mini.
arXiv:2510. 19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and solve them autonomously.
arXiv:2608.29397v1 Announce Type: new Abstract: Tool-use benchmarks generally evaluate whether an agent completes a workflow using appropriate tools and valid arguments. However, feasibility alone is...