arXiv AI

From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs

arXiv:2606. 17648v1 Announce Type: new Abstract: Standard accuracy metrics cannot explain why LLMs handle variable tracking but fail on semantically equivalent loops.

arXiv Computation and Language
Sep 25

Encoded but Not Decoded: Layer-Localized Evidence for a Three-Level Gap in LLM Syntax

The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.

By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv AI
Sep 2

Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls

The paper introduces a benchmark for evaluating large language models (LLMs) on long‑horizon state tracking by having them compute the MD5 hash through 196 dependent tool calls across 64 rounds, carrying four 32‑bit words in context. It shows that a mixture‑of‑experts LLM can maintain the full state and produce correct digests in most runs, even when all primitive tools are replaced by another LLM. The study isolates state‑tracking difficulty from instruction interpretation and identifies key factors—contextual reasoning and worker voting—that enable success.

By Dheeraj Mohandas Pai, Lu Xian