arXiv AI

Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry

arXiv:2606. 13934v1 Announce Type: new Abstract: Humans cannot always intuit what scenarios are most challenging to LLMs.

arXiv Machine Learning
Jul 24

Concept Concentration for Faithful Representation Intervention

arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.

By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv AI
Jun 6

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

arXiv:2606. 05290v1 Announce Type: cross Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture.

By Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara