Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
arXiv:2607. 25467v1 Announce Type: cross Abstract: Stateful multimodal assistants encode an image once but may answer questions about it many turns later.
arXiv:2606. 29788v1 Announce Type: new Abstract: When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success.
arXiv:2607. 25467v1 Announce Type: cross Abstract: Stateful multimodal assistants encode an image once but may answer questions about it many turns later.
arXiv:2511. 20196v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training.
arXiv:2606. 24112v1 Announce Type: new Abstract: Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors.
arXiv:2606. 27499v1 Announce Type: cross Abstract: Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environment, when an agent genuinely needs to remember what it saw rather than what it could write down.
arXiv:2608. 03791v1 Announce Type: new Abstract: Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora.
arXiv:2607. 18615v1 Announce Type: cross Abstract: Machine unlearning for vision-language models (VLMs) remains underexplored.
Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpora. Removing such knowledge is essential for building trustworthy AI systems.
arXiv:2607. 11919v1 Announce Type: cross Abstract: Human memory is reconstructive, not a faithful recording.
arXiv:2606. 10742v1 Announce Type: cross Abstract: External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences.
arXiv:2607. 07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data.
arXiv:2607. 12278v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) for computational pathology report striking zero-shot performance on whole-slide image (WSI) visual question answering (VQA) benchmarks.
arXiv:2608. 17722v1 Announce Type: cross Abstract: Vision-Language models (VLMs) achieve outstanding performance largely due to the amount of training data available on the internet.