arXiv AI

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

The article discusses how corrections made by experts to large language model (LLM) assistants often fail to persist beyond a session, framing this as an operations issue rather than a tooling one. The author, a seasoned systems engineer, maps the LLM stack onto traditional engineering components—such as frozen silicon, firmware, and persistent configuration—to highlight gaps in stochastic generation and rule retirement. From these gaps, a seven‑principle operating discipline is proposed, centered on an error loop, and illustrated with three real‑world cases, including a control that inadvertently caused the harm it was meant to prevent. The paper concludes by outlining a measurement framework and a lab study needed to validate the approach.

Hugging Face Trending Papers
Aug 19

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

The paper argues that when a human corrects an LLM assistant’s mistake, the correction often disappears after the session ends, highlighting an operations issue rather than a tooling one. Drawing on thirty years of systems engineering experience, the author maps the LLM stack onto traditional hardware and software components, identifies mismatches—such as stochastic generation and lack of a retirement stage—and proposes a seven‑principle operating discipline centered on an error loop. The paper includes three real‑world cases, one of which illustrates how a control mechanism can inadvertently cause the harm it was meant to prevent, and concludes with a suggested measurement framework and a lab study to validate the approach.

arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan
Hugging Face Trending Papers
Sep 3

Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints

The paper investigates the reliability of language‑model judges used as measurement instruments on shared endpoints. Through two preregistered audits of 52,988 requests, the authors found that repeat rankings and byte‑identical replays fell far short of required thresholds, revealing significant instability. They identify three mechanisms—label‑to‑meaning bias, candidate gaps below the noise floor, and input permutation noise—that explain the gap, and propose a snapshot‑identity ladder, design rules, and a reporting checklist to mitigate such failures.

arXiv AI
Aug 28

Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy

The paper reports a failure study of a production agentic software‑delivery platform, analyzing 147 incidents across 81 runs. It shows that the standard reliability primitives—retry, timeout, and error‑rate circuit breaking—fail in practice, leading to costly loops, false trips, and blocked work. The authors identify two cross‑cutting causes—identity adequacy and evidence adequacy—and propose seven new reliability primitives that enforce reliability at the delegation level.

By Mazhar Shaikh, Anurag Rajkumar Bombarde, Harshal Pathak