arXiv:2602. 07008v3 Announce Type: replace-cross Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence.
By Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Sanyi Zhang, Kangwei Liu, Shiming Liu, Zhangcheng Wang, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao
arXiv:2607. 23804v1 Announce Type: cross Abstract: Context attribution methods for large language models (LLMs) identify which input context contributes to the model response.
By Quoc-Huy Trinh, Lin Zhu, Sebastian Szyller
arXiv:2606. 10877v1 Announce Type: new Abstract: Occlusion-based attribution methods provide an intuitive way to estimate feature importance by perturbing input features and measuring the resulting change in model output.
By Thodoris Lymperopoulos, Ioannis Kakogeorgiou, Denia Kanellopoulou
arXiv:2604.05819v2 Announce Type: replace-cross
Abstract: Interpreting the decisions of complex computer vision models is crucial to establish trust and accountability, especially in safety-critical...
By David Schinagl, Christian Fruhwirth-Reisinger, Alexander Prutsch, Samuel Schulter, Horst Possegger
arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.
By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv:2607. 12052v1 Announce Type: cross Abstract: Synthetic image attribution aims at identifying the generator responsible for a given AI-generated image.
By Meiling Li, Pietro Bongini, Benedetta Tondi, Mauro Barni
arXiv:2606. 04928v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and data provenance.
By Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Kaan Bayraktar, Roger Wattenhofer
arXiv:2505. 03201v4 Announce Type: replace-cross Abstract: Integrated Gradients (IG) is a widely used attribution method in explainable AI, particularly in computer vision applications where reliable feature attribution is essential.
By Kien Tran Duc Tuan, Tam Nguyen Trong, Son Nguyen Hoang, Khoat Than, Anh Nguyen Duc
arXiv:2510. 12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes.
By Noor Islam S. Mohammad, Ulu\u{g} Bayaz{\i}t
The paper proposes a two‑stage framework, SL+LHF, that first learns low‑dimensional representations from noisy labeled data and then refines model alignment using human comparison feedback via a probabilistic bisection approach. It introduces the label‑noise‑to‑comparison‑accuracy (LNCA) ratio to theoretically identify when this framework outperforms pure supervised learning, showing that trading labels for comparisons reduces sample complexity when labels are scarce. Experiments on a high‑dimensional crowdfunding prediction task and an Amazon Mechanical Turk study confirm that incorporating human or large language model evaluators improves accuracy under a fixed query budget.
By Junyu Cao, Mohsen Bayati
arXiv:2606. 03885v1 Announce Type: new Abstract: Feature attribution methods explain predictions by assigning importance scores to input features.
By Kieran A. Murphy, Shameen Shrestha
arXiv:2606. 26387v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) extend large language models (LLMs) with visual perception, enabling joint reasoning over images and text.
By Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, Tianyang Wang, Hao Xu