AdaMEM: Test-Time Adaptive Memory for Language Agents
arXiv:2606. 05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions.
arXiv:2606. 11182v1 Announce Type: cross Abstract: In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams.
arXiv:2606. 05684v1 Announce Type: new Abstract: A central challenge for language agents is utilizing past experience to adapt to dynamic test-time conditions.
Combee is a new framework that scales prompt learning for self‑improving language model agents by enabling many agents to run in parallel while learning from their combined traces. It uses parallel scans, an augmented shuffle mechanism, and a dynamic batch size controller to maintain quality and reduce delay. Experiments on AppWorld, Terminal‑Bench, Formula, and FiNER show up to 17× speedup over prior methods with comparable or better accuracy at similar cost.
arXiv:2606. 01279v1 Announce Type: new Abstract: AI agents are increasingly being tasked with automating AI research itself, particularly the critical post-training phase that transforms base LLMs into aligned assistants.
arXiv:2609.23600v1 Announce Type: new Abstract: Prompt learning efficiently adapts vision-language models (VLMs) to downstream tasks, but gains on seen classes often come at the expense of generaliza...
arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.
arXiv:2608. 15639v1 Announce Type: cross Abstract: \textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients.
arXiv:2603. 20895v3 Announce Type: replace-cross Abstract: Existing routers rely on semantic query features or handcrafted features, which often fail to capture model-specific failures or intrinsic task difficulty.
arXiv:2606. 01770v1 Announce Type: cross Abstract: Auto-harness systems such as A-Evolve, GEPA, and Meta-Harness improve LLM agents by optimizing prompts, skills, tools, memories, and supporting infrastructure from execution feedback, but they are typically evaluated on fixed offline benchmarks.
The paper introduces Test-Time Environment Decomposition (TTED), a label‑free learning method that allows large language model agents to break down complex web environment observations into simpler sub‑modules during inference. By learning from experience within these sub‑environments, agents can compose the gained knowledge to improve performance in the full environment. Experiments on synthetic and realistic benchmarks show that this approach enhances compositional generalization and boosts real‑world web automation tasks.
arXiv:2603. 17216v2 Announce Type: replace Abstract: With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal.
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification. Yet current evaluations do not isolate this form of transfer.
arXiv:2606. 09430v1 Announce Type: cross Abstract: Online task-free continual learning (TFCL) requires intelligent agents to sequentially accumulate knowledge from an unbounded, non-stationary data stream under strict single-pass constraints and without any explicit task identifiers.