arXiv:2606. 30587v1 Announce Type: cross Abstract: Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection.
By Asif Shahriar, Hongyu Cai, Hadjer Benkraouda, Gang Wang, Z. Berkay Celik
The paper presents a taxonomy-driven framework for identifying, categorizing, and explaining bias in AI-generated Python code. By extending an existing dataset and manually annotating bias categories and justifications, the authors evaluate both proprietary and open-source large language models (LLMs) for automated bias detection and explanation. Results show that models such as Gemini and Qwen3-coder achieve high classification accuracy and produce justification and code identification similarities that closely match human-authored reasoning.
By Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
arXiv:2509. 20491v3 Announce Type: replace-cross Abstract: Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code.
By Brahim Mahmoudi, Naouel Moha, Quentin Sti\'evenart, Florent Avellaneda
arXiv:2609.22119v1 Announce Type: cross
Abstract: Evaluation awareness poses an unprecedented threat to model evaluation, but the mechanisms by which models detect it remain unknown. This study focus...
By Navraj Singh, Maheep Chaudhary
arXiv:2509. 14335v2 Announce Type: replace-cross Abstract: Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence.
By Xinran Zheng, Xingzhi Qian, Yiling He, Shuo Yang, Lorenzo Cavallaro
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.