← Back to all news
arXiv Machine Learning September 15, 2026 By Aditya Karnam Gururaj Rao, Arjun Jaggi

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • agents
  • nlp

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 10

Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models

arXiv:2607. 08665v1 Announce Type: new Abstract: Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle.

By Teng-Ruei Chen
llmsbenchmarks
More like this →
arXiv AI
Sep 23

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

arXiv:2609.23790v1 Announce Type: new Abstract: Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected toke...

By Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick
llmsagentsbenchmarks
More like this →
arXiv Machine Learning
Jun 25

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.

By Alex Kwon
llmsbenchmarks
More like this →
arXiv AI
Jul 15

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

arXiv:2607. 12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task.

By Brenda Lelis, Rodrigo Cabral-Carvalho
llmsagentsbenchmarks
More like this →
arXiv AI
Jul 22

Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary

arXiv:2607. 18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes?

By Jan Kirin
llmsfine-tuning
More like this →
arXiv AI
Aug 17

MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends

arXiv:2608. 13883v1 Announce Type: new Abstract: Most agent-memory benchmarks test post-hoc recall, whereas MemoryArena evaluates whether memory supports interdependent, multi-session task completion.

By Chaoqun Zhan, Qiang Zhou, Guannan Li, Zhenqiang Huang, Qianjin Wang
llmsragagentsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea