arXiv:2607. 27281v1 Announce Type: new Abstract: A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing.
By Lei Dong
The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.
By Chencheng Zhu
The paper introduces GRADE, a graph-based representation of large language model (LLM) agent executions that captures both execution steps and their dependencies. By adding graded dependency edges—observed, declared, or inferred—to the trace, the authors evaluate how this dependency layer affects failure prediction across six corpora involving tool use, coding, and web tasks. Experiments show that the dependency block can improve prediction in some settings, but its effectiveness varies with the evaluation probe and corpus, and controlled experiments demonstrate that the observed structure is not merely a degree-matched artifact.
By Yue Zhao
arXiv:2607. 29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions.
By Xiyang Zhang, Hongzhi Wang, Yuanhe Tian
arXiv:2608.30427v1 Announce Type: cross
Abstract: Speculative decoding speeds up generation with an efficient draft model (drafter) that proposes tokens for a target model to verify in one pass, pres...
By Ephrem Wu
arXiv:2609.25052v1 Announce Type: new
Abstract: "An agent that writes its conclusions into a store it later retrieves from closes a loop usually reported as one-way contamination. Taking the loop to...
By Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong
arXiv:2609.23886v1 Announce Type: new
Abstract: Software delegates more of its branches to models every year: which queue a ticket enters, whether a command is safe to run, whether a claim clears wit...
By Zehua Cheng, Wei Dai, Jiahao Sun
arXiv:2608. 11318v1 Announce Type: cross Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent.
By Yi Liu
arXiv:2607. 01893v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, committing only the longest accepted prefix.
By Tianjian Yang, Meng Li
The paper investigates how inherited state affects sub-agent performance in multi-agent frameworks, comparing three inheritance policies—Reset, Selective, and Full—across a ladder of Qwen3 models. It finds that reliance on stale state decreases with model capability, but a mid-capability model (Qwen3‑1.7B) exhibits a statistically significant local minimum of net harm, defining a "danger band." Selective handoff consistently improves accuracy over Full, especially within the danger band, while a fixed-threshold router fails on other datasets.
By Jundong Hu, Shekar Ramachandran
The paper introduces instruction duplication, a simple inference‑time control that repeats the procedural instruction without retraining or decoding changes. Across seven instruction‑tuned models and 16,800 scheduled generations, duplicating the instruction improves deterministic All‑8 diagnostic‑response success from 90.22% to 93.17% and reduces failures by 30.2%. In downstream Answer Engineering scenarios, duplication further boosts success rates, demonstrating its practical impact on systems that rely on the generated trajectory.
By Victor Lavrenko (PeaceTech VC, Israel)
arXiv:2609.11987v1 Announce Type: cross
Abstract: An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous softwar...
By Mohsen Arjmandi