arXiv AI

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

arXiv:2608. 10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world.

arXiv AI
Jun 19

How Transparent is DiffusionGemma?

arXiv:2606. 20560v1 Announce Type: cross Abstract: LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors.

By Joshua Engels, Callum McDougall, Bilal Chughtai, Janos Kramar, Senthoran Rajamanoharan, Cindy Wu, Arthur Conmy, Asic Q Chen, Jean Tarbouriech, Min Ma, Brendan O'Donoghue, Jo\~ao Gabriel Lopes de Oliveira, Rohin Shah, Neel Nanda
arXiv AI
Sep 15

Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference

The paper introduces self‑orchestrating language models that annotate semantic dependence—identifying which tokens rely on others—to guide efficient inference. By leveraging these annotations, the authors design runtimes that parallelize autoregressive decoding, evict intermediate context, or determine denoising orders, achieving Pareto‑optimal quality‑efficiency trade‑offs. Three systems—PASTA, TIP, and Planned Diffusion—demonstrate these techniques for parallel decoding, memory‑efficient reasoning, and efficient discrete diffusion, respectively.

By Tian Jin
arXiv AI
Jun 15

Residual Context Diffusion Language Models

arXiv:2601. 22954v2 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can decode multiple tokens in parallel.

By Yuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi, Coleman Hooper, Jintao Zhang, Aditya Tomar, Michael W. Mahoney, Sewon Min, Mehrdad Farajtabar, Kurt Keutzer, Amir Gholami, Chenfeng Xu
arXiv AI
Sep 25

When Can Agents Forget Their Reasoning? ICLR for Long-Horizon Agent Context Compression

Long‑horizon language model agents accumulate reasoning history, which inflates context length and inference cost. The paper introduces Interaction Aware Compression for Long Horizon Reasoning (ICLR), a training‑free online method that ranks and removes reasoning blocks based on frozen proxy entropy while preserving actions, tool calls, and observations. On 260 WorkBuddyBench tasks, ICLR raises average reward from 0.699 to 0.718 and cuts input, output, and cache read tokens by 25.5%, 14.4%, and 33.3% respectively, while analyses show that historical reasoning becomes replaceable once task‑relevant state is externalized.

By Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv Computation and Language
Aug 24

Self-Speculation for Faster Reasoning Models

arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performanc...

By Ravisri Valluri, Tung Nguyen, Aditya Grover