arXiv AI

Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts

arXiv:2607. 10411v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics.

arXiv AI
6d ago

A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code

The paper presents a taxonomy-driven framework for identifying, categorizing, and explaining bias in AI-generated Python code. By extending an existing dataset and manually annotating bias categories and justifications, the authors evaluate both proprietary and open-source large language models (LLMs) for automated bias detection and explanation. Results show that models such as Gemini and Qwen3-coder achieve high classification accuracy and produce justification and code identification similarities that closely match human-authored reasoning.

By Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez
arXiv Machine Learning
Jun 5

Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation

arXiv:2606. 05792v1 Announce Type: cross Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language still requires time and expertise, which limits adoption.

By Arslan Bisharat, Brian Ortiz, Eric Spencer, Khushboo Bhadauria, TaiNing Wang, George K. Thiruvathukal, Konstantin Laufer, Mohammed Abuhamad