The article discusses how corrections made by experts to large language model (LLM) assistants often fail to persist beyond a session, framing this as an operations issue rather than a tooling one. The author, a seasoned systems engineer, maps the LLM stack onto traditional engineering components—such as frozen silicon, firmware, and persistent configuration—to highlight gaps in stochastic generation and rule retirement. From these gaps, a seven‑principle operating discipline is proposed, centered on an error loop, and illustrated with three real‑world cases, including a control that inadvertently caused the harm it was meant to prevent. The paper concludes by outlining a measurement framework and a lab study needed to validate the approach.
By George Andrikopoulos
arXiv:2607. 11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work.
By Aritra Mazumder, Nusrat jahan Lia
Everyone is talking about loop engineering, but most discussions assume an LLM sits at the center of the loop. I wanted to isolate the architecture itself.
By Emmimal P Alexander
arXiv:2607. 13071v1 Announce Type: cross Abstract: Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth.
By Hiroki Tamba
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
By Esmail Gumaan
arXiv:2606. 01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory, and recovery.
By Rahul Suresh Babu, Adarsh Agrawal