arXiv AI

Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor

arXiv:2507. 15903v2 Announce Type: replace-cross Abstract: Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilitate AI deployment.

arXiv AI
Aug 28

Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention

The paper surveys hallucinations in large language models (LLMs) through a lifecycle lens, covering causes, detection, mitigation, and prevention. It categorizes hallucinations into data‑related, training‑related, and inference‑related stages, aligning each with specific interventions. The authors also review benchmark datasets and propose a standardized framework to diagnose and address hallucinations for safer, more reliable LLMs.

By Naveen Lamba, Sanju Tiwari, Manas Gaur
arXiv AI
Sep 24

Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory

The paper introduces ShadowMem, a defensive framework that protects large language model agents from long-horizon threats by maintaining a dedicated safety-focused memory. Inspired by the shadow stack concept, ShadowMem stores safety-critical context throughout an agent’s execution and uses this shadow memory to evaluate the risk of upcoming actions before they are carried out. Experiments show that ShadowMem outperforms existing defenses in detection accuracy, detects most attacks early, and adds minimal overhead to agent performance.

By Yuhui Wang, Tanqiu Jiang, Jiacheng Liang, Charles Fleming, Ting Wang