Towards Data Science By Miodrag Cekikj

Designing a Persistent Knowledge Layer That Refuses to Guess

Read the original on Towards Data Science →

RAG Retrieves, It Never Remembers. A vendor-neutral blueprint for applications that accumulate understanding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Towards Data Science.

arXiv Machine Learning
Sep 15

Realistic Continual Learning Approach using Pre-trained Models

arXiv:2404.07729v2 Announce Type: replace Abstract: Continual learning (CL) evaluates adaptability in learning solutions to retain knowledge. Our research addresses the challenge of catastrophic forg...

By Nadia Nasri, Carlos Guti\'errez-\'Alvarez, Sergio Lafuente-Arroyo, Saturnino Maldonado-Basc\'on, Roberto J. L\'opez-Sastre
arXiv AI
Sep 15

Homeostatic Continual Learning

The paper introduces a new approach called Homeostatic Continual Learning, designed to allow an AI agent to learn continuously in a changing environment without catastrophic forgetting. The method identifies outliers in environmental data when the agent’s output deviates, enabling the agent to incrementally refine its model and policy across increasingly diverse contexts. The authors also propose extending the method to build a world model that factorizes objects into features, abstracts them into comparable concept instances, and maps concepts to intents via features, while outlining necessary future work and broader AI connections.

By Yue Jin
arXiv AI
Sep 15

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

K-Bench is a new benchmark designed to evaluate large language model (LLM) unlearning when the models are deployed as agents. Unlike previous benchmarks that only inspect the final answer, K-Bench examines all six channels of a ReAct agent—including chain-of-thought, tool calls, tool observations, and elicited summaries—to determine if a secret is leaked. The benchmark measures leakage for secrets placed in the model weights, prompt, or retrieval store, and finds that many existing unlearning methods fail to prevent leaks in deployed agents, especially when secrets reside in the prompt or retrieval store.

By Guangsheng Yu, Yanna Jiang, Qin Wang, Baihe Ma, Xu Wang