arXiv Machine Learning

Concept Labels Are Not Enough: Rethinking Concept Bottleneck Models through Representation Integrity

arXiv:2510. 15770v4 Announce Type: replace-cross Abstract: Although deep neural networks achieve strong predictive performance, their internal reasoning often remains difficult to inspect and control.

arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han