Towards Data Science

Can an LLM Forget the Right Things?

The article discusses a specialized LLM inference runtime designed for real-time applications, such as a 33 ms robot control cycle. Unlike typical runtimes that ignore physical deadlines, this system refuses new requests when the deadline is at risk, evicts key‑value cache entries based on meaning rather than age, and is implemented entirely in hand‑written CUDA without relying on cuBLAS or libtorch.

Simon Willison
Sep 18

Note on 18th September 2026

Simon Willison reflects on his current disinterest in large language models (LLMs), comparing it to a geneticist dismissing the newly opened Jurassic Park. He emphasizes that this stance feels odd given the excitement surrounding LLMs. The note highlights his personal stance on AI and generative‑AI topics.

Towards Data Science
Aug 11

Can a Local LLM Run My AI Assistant?

I replayed the same 27 real production tasks through two local models, one hardware upgrade apart, to find out what it actually takes to replace Claude as the brain behind a 90-tool personal agent. The post Can a Local LLM Run My AI Assistant?

By Arsen Apostolov
arXiv AI
Sep 2

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

The paper introduces the Prospective Intention Store (PIS), a method that places lifecycle logic in code and confines language tasks to a typed action space, enabling small models to perform prospective memory tasks more effectively. Using PIS, a small model (DeepSeek-Chat) achieves 82.9% Set‑F1 on PM‑Bench, surpassing the previous best of 65.1%. On Gemma‑E2B, PIS boosts Set‑F1 from 4.2% (without a store) to 66.2%, and reaches 70.1% Set‑F1, outperforming retrospective memory approaches that max out at 54.4%.

By Jinqing Zhao, Chengcan Wu
arXiv AI
4d ago

Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents

The paper introduces the Distributed‑Evidence Paradox, where long‑running LLM agents compress past interactions into persistent memories that may not be fully supported by the interaction history. It defines three key requirements—evidence scope, compositional validity, and admission reliability—and proposes DerivAudit, a framework that checks whether a memory is truly supported by the available history. Experiments on two memory corpora show that expanding the evidence base can recover support for many memories, yet many remain unsupported, and broader evidence alone does not guarantee reliable admission.

By Hongjun Liu, Chen Zhao
Hugging Face Trending Papers
Aug 5

Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems

LLM agents that persist across sessions accumulate stored memories whose validity varies enormously by content type, yet existing memory architectures treat all memories as equally persistent and systematically contaminate retrieved context with outdated facts. We show that per-memory, type-conditioned temporal decay, a property of western scrub jay episodic memory, can be operationalized as an auto-classified coefficient $π_i$ in an external LLM-agent memory store, yielding ScrubJay-MEM: each memory is encoded as a jointly-bound What--Where--When tuple with an estimated perishability $π_i$ and utility horizon $τ_i$, retrieved by query-adaptive scoring, and revised retroactively at $O(1)$ LLM calls per update.

arXiv Computation and Language
Sep 1

Hindsight Memory-PRM: Supervising Memory Management with Auditable Hindsight Credit

arXiv:2608.29605v1 Announce Type: new Abstract: Memory operations of long-horizon LLM agents are hard to supervise: an operation's value is unobservable when it is taken. But they are special -- they...

By Haoxuan Jia, Yang Liu, Yingguang Yang, Yancheng Chen, Chongyang Zhang, Hao Zheng, Qian Li, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Hao Peng, Junyu Lu, Du Cheng, Philip S. Yu, Bin Chong