arXiv AI By Francesco Sovrano

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP

Read the original on arXiv AI →

arXiv:2505. 11189v3 Announce Type: replace Abstract: Large language models (LLMs) can amplify misinformation, undermining societal goals such as the UN SDGs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.