iFlip is an iterative refinement method for generating counterfactual examples using large language models. It incorporates three feedback types—model confidence, feature attribution, and natural language—to guide successive edits. Experiments show iFlip outperforms five state‑of‑the‑art baselines, achieving a 57.8% higher validity rate and improving model performance through counterfactual data augmentation.
By Yilong Wang, Qianli Wang, Nils Feldhus
The paper introduces Causal-Counterfactual RAG, a new framework that augments Retrieval-Augmented Generation with explicit causal graphs and counterfactual reasoning. By incorporating cause‑effect relationships into retrieval and evaluating both direct causal evidence and counterfactual scenarios, the approach aims to produce more robust, accurate, and interpretable answers. This method seeks to maintain contextual coherence, reduce hallucinations, and improve reasoning fidelity compared to traditional RAG systems.
By Harshad Khadilkar, Abhay Gupta
arXiv:2606.03695v2 Announce Type: replace
Abstract: As language models are increasingly deployed in real-world applications, the ability to erase specific knowledge from them becomes critical for saf...
By Clara Haya Suslik, Or Shafran, Mor Geva
arXiv:2606. 06320v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities.
By Gizem Y\"uce, Giorgos Nikolaou, Nicolas Flammarion
arXiv:2609.38929v1 Announce Type: new
Abstract: Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topica...
By Pinaki Mohanty, Haoran Tang, Maggie Makar, Rajiv Khanna
GRACE is a new framework for concept erasure in text-to-image diffusion models that uses a semantically weighted sensitive subspace to guide localized interventions and lightweight subspace-constrained adapters to avoid global semantic disruption. It replaces manual counterfactual prompts with an automatically decoupled safe-anchor mechanism and controls intervention strength through an energy-driven dynamic gating system. Experiments show GRACE improves NSFW reduction by 17.86% over five state-of-the-art methods while also reducing target CLIP Score and FID, indicating stronger concept suppression with better preservation of generative quality.
By Qinghui Gong, Yihuai Liang, Yuanlun Xie, Deepak Kumar Jain, Vitomir \v{S}truc, Zhengchun Zhou