Towards Data Science

Context Engineering Isn’t Enough — A Loop Engineering Experiment With No LLM Inside the Loop

Everyone is talking about loop engineering, but most discussions assume an LLM sits at the center of the loop. I wanted to isolate the architecture itself.

Hugging Face Trending Papers
Aug 19

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

The paper argues that when a human corrects an LLM assistant’s mistake, the correction often disappears after the session ends, highlighting an operations issue rather than a tooling one. Drawing on thirty years of systems engineering experience, the author maps the LLM stack onto traditional hardware and software components, identifies mismatches—such as stochastic generation and lack of a retirement stage—and proposes a seven‑principle operating discipline centered on an error loop. The paper includes three real‑world cases, one of which illustrates how a control mechanism can inadvertently cause the harm it was meant to prevent, and concludes with a suggested measurement framework and a lab study to validate the approach.

arXiv AI
Aug 20

Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering

The article discusses how corrections made by experts to large language model (LLM) assistants often fail to persist beyond a session, framing this as an operations issue rather than a tooling one. The author, a seasoned systems engineer, maps the LLM stack onto traditional engineering components—such as frozen silicon, firmware, and persistent configuration—to highlight gaps in stochastic generation and rule retirement. From these gaps, a seven‑principle operating discipline is proposed, centered on an error loop, and illustrated with three real‑world cases, including a control that inadvertently caused the harm it was meant to prevent. The paper concludes by outlining a measurement framework and a lab study needed to validate the approach.

By George Andrikopoulos
arXiv AI
Sep 23

World State Generator

arXiv:2609.24744v1 Announce Type: new Abstract: Language agents solve complex tasks through plans and actions. A single step the world refuses puts the goal out of reach, and what the agent does next...

By Sungheon Jeong, Sanggeon Yun, Ryozo Masukawa, Haleh Alimohamadi, Mahdi Imani, Mohsen Imani
arXiv AI
Sep 2

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

The paper introduces a benchmark for evaluating large language models (LLMs) on long‑horizon state tracking by having them compute the MD5 hash through 196 dependent tool calls across 64 rounds, carrying four 32‑bit words in context. It shows that a mixture‑of‑experts LLM can maintain the full state and produce correct digests in most runs, even when all primitive tools are replaced by another LLM. The study isolates state‑tracking difficulty from instruction interpretation and identifies key factors—contextual reasoning and worker voting—that enable success.

By Dheeraj Mohandas Pai, Lu Xian
arXiv AI
Jul 24

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya