arXiv:2609.36931v1 Announce Type: new
Abstract: Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked...
By Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, Katharina von der Wense
Simon Willison reflects on his current disinterest in large language models (LLMs), comparing it to a geneticist dismissing the newly opened Jurassic Park. He emphasizes that this stance feels odd given the excitement surrounding LLMs. The note highlights his personal stance on AI and generative‑AI topics.
Most AI memory systems keep the newest information—not the most important. Here's how I used the Ebbinghaus forgetting curve to build a better memory engine for LLMs.
By Emmimal P Alexander
arXiv:2607. 22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning.
By Yan Zhang, Shibo Li
I replayed the same 27 real production tasks through two local models, one hardware upgrade apart, to find out what it actually takes to replace Claude as the brain behind a 90-tool personal agent. The post Can a Local LLM Run My AI Assistant?
By Arsen Apostolov
The paper introduces the Prospective Intention Store (PIS), a method that places lifecycle logic in code and confines language tasks to a typed action space, enabling small models to perform prospective memory tasks more effectively. Using PIS, a small model (DeepSeek-Chat) achieves 82.9% Set‑F1 on PM‑Bench, surpassing the previous best of 65.1%. On Gemma‑E2B, PIS boosts Set‑F1 from 4.2% (without a store) to 66.2%, and reaches 70.1% Set‑F1, outperforming retrospective memory approaches that max out at 54.4%.
By Jinqing Zhao, Chengcan Wu