arXiv AI By Jipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song

Can Computation from Earlier Problems Help LLMs Solve New Ones?

Read the original on arXiv AI →

The paper investigates whether computation from earlier problems can aid large language models (LLMs) in solving subsequent ones. Preliminary experiments show that retained conversation history can both improve and degrade later-turn accuracy. To address this, the authors propose STAIR, a lightweight module that stores key-value pairs from previous responses and redirects new queries to this bank, improving later-turn accuracy by up to 11.67 percentage points across several benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

When Context Changes: Understanding Update Failures in LLMs

The paper introduces the concept of stale binding, where large language models (LLMs) answer with outdated values despite having newer information in context. It presents Controlled In-Context Memory (CICM), a benchmark to track and test the use of updated information in conversations and agent logs. Experiments on open‑source models reveal an attention drift mechanism that favors old values, and the authors propose a simple attention‑redirecting intervention that largely corrects these errors without harming correct answers.

By Junyu Guo, Yuchen Fang, Shangding Gu, Costas Spanos, James Demmel, Javad Lavaei
arXiv AI
4d ago

Mnemon: Raw Records, Fast Judgments, Slow Thoughts

Mnemon is a memory agent that stores conversations as raw, dated records and uses a fast System 1 decision model (Jev) to quickly judge the relevance of records, while a slow System 2 LLM plans searches and composes answers. The agent consolidates records into topic timelines and value histories in the background, enabling efficient retrieval without rewriting conversations into structured formats. Experiments show Mnemon achieving high scores on LoCoMo and LongMemEval‑S with low context length and cost, and Jev outperforming LLMs in evidence separation and speed.

By Guangren Wang