arXiv:2605.01913v2 Announce Type: replace-cross
Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vul...
By Sadia Asif, Mohammad Mohammadi Amiri
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2511. 05865v3 Announce Type: replace-cross Abstract: Recent advancements in large-scale generative models have enabled the creation of high-quality images and videos, but have also raised significant safety concerns regarding the generation of unsafe content.
By Viet Nguyen, Vishal M. Patel
arXiv:2512. 20638v2 Announce Type: replace-cross Abstract: The evaluation of large language models relies heavily on standardized benchmarks.
By Maty Bohacek, Nino Scherrer, Nicholas Dufour, Thomas Leung, Christoph Bregler, Stephanie C. Y. Chan
arXiv:2608. 04286v1 Announce Type: cross Abstract: Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG).
By Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
arXiv:2606. 30498v1 Announce Type: cross Abstract: Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color.
By Laines Schmalwasser, Jan Blunk, Niklas Penzel, Julia Niebling, Joachim Denzler
arXiv:2511. 19636v2 Announce Type: replace-cross Abstract: In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fundamentally different internal logic.
By Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong
arXiv:2505. 18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors.
By Hongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu, Chaowei Xiao, Kun Zhang, Bo Han
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
arXiv:2608. 20338v1 Announce Type: new Abstract: Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this capability completely.
By Sahil Kale, Ian Harris
arXiv:2509. 22015v2 Announce Type: replace Abstract: Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, providing a powerful lens for passive feature discovery.
By Jianrong Ding, Muxi Chen, Chenchen Zhao, Qiang Xu
arXiv:2607. 08317v1 Announce Type: new Abstract: Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a dog with five legs.
By Matteo Santelmo, Xiuying Wei, Israa Fakih, Felix Bauer, Juan Garcia Giraldo, Chengkun Li, Etienne Bamas, Emmanuel Abb\'e