Towards Data Science By Anubhab Banerjee

Can an LLM Forget the Right Things?

Read the original on Towards Data Science →

The article discusses a specialized LLM inference runtime designed for real-time applications, such as a 33 ms robot control cycle. Unlike typical runtimes that ignore physical deadlines, this system refuses new requests when the deadline is at risk, evicts key‑value cache entries based on meaning rather than age, and is implemented entirely in hand‑written CUDA without relying on cuBLAS or libtorch.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

Simon Willison
Sep 18

Note on 18th September 2026

Simon Willison reflects on his current disinterest in large language models (LLMs), comparing it to a geneticist dismissing the newly opened Jurassic Park. He emphasizes that this stance feels odd given the excitement surrounding LLMs. The note highlights his personal stance on AI and generative‑AI topics.

Towards Data Science
Aug 11

Can a Local LLM Run My AI Assistant?

I replayed the same 27 real production tasks through two local models, one hardware upgrade apart, to find out what it actually takes to replace Claude as the brain behind a 90-tool personal agent. The post Can a Local LLM Run My AI Assistant?

By Arsen Apostolov
arXiv AI
Sep 2

Making Prospective Memory SLM-Shaped: Typed Intention Stores for Small-Model Agents

The paper introduces the Prospective Intention Store (PIS), a method that places lifecycle logic in code and confines language tasks to a typed action space, enabling small models to perform prospective memory tasks more effectively. Using PIS, a small model (DeepSeek-Chat) achieves 82.9% Set‑F1 on PM‑Bench, surpassing the previous best of 65.1%. On Gemma‑E2B, PIS boosts Set‑F1 from 4.2% (without a store) to 66.2%, and reaches 70.1% Set‑F1, outperforming retrospective memory approaches that max out at 54.4%.

By Jinqing Zhao, Chengcan Wu