← Back to all news
arXiv Computation and Language August 25, 2026 By Ayush Gupta, Hima Varshini Surisetty, Sreevidya Bollineni, Varad Ingale, Tuhina Tripathi, Abhishek Lalwani, Somya Chatterjee, Sadid Hasan

Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • fine-tuning
  • multimodal
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 6

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

arXiv:2606. 06286v1 Announce Type: cross Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather than whether they do so under ordinary use.

By Gianluca Barmina, Peter Schneider-Kamp, Lukas Galke Poech
llmssafety
More like this →
arXiv Machine Learning
Jun 16

Auditing Machine Unlearning: A Systematic Research on Whether Models Truly Forget

arXiv:2606. 16110v1 Announce Type: new Abstract: Machine unlearning has been extensively studied in response to growing privacy concerns and regulatory requirements.

By Dayong Ye, Tianqing Zhu, Ruiding Huang, Xinbo Fu, Jiayang Li, Bo Liu, Huan Huo, Wanlei Zhou
llmsfine-tuningsafety
More like this →
arXiv Machine Learning
Jun 2

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs

arXiv:2506. 14003v5 Announce Type: replace Abstract: Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowledge from a trained model, while maintaining its performance on standard tasks.

By Yiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu, Sijia Liu
llmssafety
More like this →
arXiv Machine Learning
Jun 18

RUB: Evaluating Residual Knowledge in Unlearned Models

arXiv:2504. 14798v2 Announce Type: replace Abstract: Machine Unlearning (MUL) has emerged as a key mechanism for privacy protection and content regulation, yet current techniques often fail to guarantee the complete removal of sensitive information.

By Hao Xuan, Xingyu Li
diffusionbenchmarkssafety
More like this →
arXiv AI
Jun 15

Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning

arXiv:2606. 14078v1 Announce Type: cross Abstract: Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks.

By Zhenqian Zhu, Yamin Hu, Yujiang Liu, Luping Wei, Wenbo Hou, Bin Li, Haodong Li, Wenjian Luo
multimodalsafety
More like this →
arXiv AI
Jul 3

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

arXiv:2607. 02513v1 Announce Type: cross Abstract: LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods.

By Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea