The paper investigates how training data attribution (TDA) can be used to influence large language models (LLMs). It compares two methods—reweighting and influence-guided response rewriting—on examples selected by influence functions. Rewriting, which replaces responses while keeping instructions fixed, yields stronger, more persistent, and bidirectional behavioral changes than reweighting, suggesting that the intervention value of influential samples is better realized through rewriting.
By Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
arXiv:2606. 18307v1 Announce Type: cross Abstract: Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs).
By Zefan Wang, Lincheng Li, Tianyu Yu, Yuan Yao
arXiv:2402. 08922v3 Announce Type: replace Abstract: Large-scale black-box models have become ubiquitous across numerous applications.
By Myeongseob Ko, Feiyang Kang, Weiyan Shi, Ming Jin, Zhou Yu, Ruoxi Jia
arXiv:2607. 22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome.
By Oleksii Furman, {\L}ukasz Lenkiewicz, Marcel Musia{\l}ek, Maciej Zi\k{e}ba
arXiv:2606. 05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions.
By Rohan N. Pradhan, Steve Goley
arXiv:2609.37616v1 Announce Type: new
Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on th...
By Abhinav Rajeev Kumar (Lossfunk), Paras Chopra (Lossfunk)
arXiv:2608.30956v1 Announce Type: cross
Abstract: Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input...
By Mattia Cerrato, Otto Sahlgren, Xenia Heilmann
arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.
By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinforcement learning. This separation is incomplete: when examples are scored and kept online during fine-tuning, the choice of which data to train on already changes the model's behavioral preferences.
arXiv:2603. 20775v2 Announce Type: replace Abstract: In personalized marketing, uplift models estimate the incremental effect of an intervention by modeling how customer behavior would change under alternative treatments using counterfactual analysis.
By Yuxuan Yang, Dugang Liu, Yiyan Huang
The paper introduces PUID, a Personalized Unobserved-Confounding-aware Interaction Deconfounder, designed to mitigate hidden confounding in recommender systems without relying on costly randomized controlled trials. PUID estimates user-item level sensitivity bounds using an entropy-based method that gauges the strength of hidden confounding from the mutual information between observed features and exposure status. An adversarial optimization strategy and a benchmark-guided variant (BPUID) further enhance robustness and predictive accuracy, and experiments on three real-world datasets show consistent outperformance over state-of-the-art baselines.
By Zongyu Li
arXiv:2607. 12985v1 Announce Type: new Abstract: Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged.
By Sen Yang, Yuen-Hei Yeung