Hugging Face Trending Papers

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

Read the original on Hugging Face Trending Papers →

The paper argues that when a human corrects an LLM assistant’s mistake, the correction often disappears after the session ends, highlighting an operations issue rather than a tooling one. Drawing on thirty years of systems engineering experience, the author maps the LLM stack onto traditional hardware and software components, identifies mismatches—such as stochastic generation and lack of a retirement stage—and proposes a seven‑principle operating discipline centered on an error loop. The paper includes three real‑world cases, one of which illustrates how a control mechanism can inadvertently cause the harm it was meant to prevent, and concludes with a suggested measurement framework and a lab study to validate the approach.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 20

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

The article discusses how corrections made by experts to large language model (LLM) assistants often fail to persist beyond a session, framing this as an operations issue rather than a tooling one. The author, a seasoned systems engineer, maps the LLM stack onto traditional engineering components—such as frozen silicon, firmware, and persistent configuration—to highlight gaps in stochastic generation and rule retirement. From these gaps, a seven‑principle operating discipline is proposed, centered on an error loop, and illustrated with three real‑world cases, including a control that inadvertently caused the harm it was meant to prevent. The paper concludes by outlining a measurement framework and a lab study needed to validate the approach.

By George Andrikopoulos
arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan