The paper argues that when a human corrects an LLM assistant’s mistake, the correction often disappears after the session ends, highlighting an operations issue rather than a tooling one. Drawing on thirty years of systems engineering experience, the author maps the LLM stack onto traditional hardware and software components, identifies mismatches—such as stochastic generation and lack of a retirement stage—and proposes a seven‑principle operating discipline centered on an error loop. The paper includes three real‑world cases, one of which illustrates how a control mechanism can inadvertently cause the harm it was meant to prevent, and concludes with a suggested measurement framework and a lab study to validate the approach.
The article discusses how corrections made by experts to large language model (LLM) assistants often fail to persist beyond a session, framing this as an operations issue rather than a tooling one. The author, a seasoned systems engineer, maps the LLM stack onto traditional engineering components—such as frozen silicon, firmware, and persistent configuration—to highlight gaps in stochastic generation and rule retirement. From these gaps, a seven‑principle operating discipline is proposed, centered on an error loop, and illustrated with three real‑world cases, including a control that inadvertently caused the harm it was meant to prevent. The paper concludes by outlining a measurement framework and a lab study needed to validate the approach.
By George Andrikopoulos
arXiv:2607. 25152v1 Announce Type: new Abstract: Long-running autonomous agents plan, act, and judge their own completion without human intervention.
By Hyundoo Park, Byungho Choi
arXiv:2607. 07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully.
By Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
arXiv:2606. 28471v1 Announce Type: new Abstract: Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score.
By Zhixuan Li, Jiangan Yuan, Han Xu
arXiv:2607. 17240v1 Announce Type: new Abstract: When does a committed intermediate stage in an LLM reasoning pipeline earn its cost?
By Honglin Li (ShanghaiTech University)
arXiv:2609.24744v1 Announce Type: new
Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next...
By Sungheon Jeong, Sanggeon Yun, Ryozo Masukawa, Haleh Alimohamadi, Mahdi Imani, Mohsen Imani
arXiv:2606. 17648v1 Announce Type: new Abstract: Standard accuracy metrics cannot explain why LLMs handle variable tracking but fail on semantically equivalent loops.
By Siyue Chen, Yifu Guo, Yuquan Lu, Zishan Xu, Jiaye Lin, Jianbo Lin, Siyu Zhang, Cheng Yang, Junxin Li, Yujia Li, Yu Huo, Ruixuan Wang
arXiv:2607. 00871v1 Announce Type: new Abstract: Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis space are produced by the policy being updated.
By Biswa Sengupta
arXiv:2606. 05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures.
By Dongsheng Zhu, Xuchen Ma, Yucheng Shen, Xiang Li, Yukun Zhao, Shuaiqiang Wang, Lingyong Yan, Dawei Yin
The paper introduces a benchmark for evaluating large language models (LLMs) on long‑horizon state tracking by having them compute the MD5 hash through 196 dependent tool calls across 64 rounds, carrying four 32‑bit words in context. It shows that a mixture‑of‑experts LLM can maintain the full state and produce correct digests in most runs, even when all primitive tools are replaced by another LLM. The study isolates state‑tracking difficulty from instruction interpretation and identifies key factors—contextual reasoning and worker voting—that enable success.
By Dheeraj Mohandas Pai, Lu Xian
arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.
By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya