arXiv Machine Learning
Sep 22

Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents

The paper audits LLM-based explainers attached to an Active Inference agent that manages German grid demand, testing three large‑language‑model backends (GPT‑4o, Claude‑3‑Opus, Gemini). By injecting corrupted observations and attacker‑controlled text, the study finds that the explainers fail to flag errors, produce fluent but incorrect rationalizations for wrong actions, and can be steered to exfiltrate data. The authors propose mitigations but do not evaluate them, emphasizing that explanations are never verified for truth before operators rely on them.

By Param Raval, Rohit Shenoy, Archana Vaidheeswaran