arXiv AI

Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation

arXiv:2608. 11513v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code.

arXiv Computation and Language
Sep 22

When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation

arXiv:2609.22102v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts...

By Anubhav Gupta, Mayara Costa Figueiredo, Leticia Santos Machado, Tanner Wright, Ivan Beschastnikh, Cleidson R. B. de Souza, Gema Rodr\'iguez-P\'erez
arXiv AI
Sep 10

Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks

The study investigates how users interact with ChatGPT for code generation beyond simple function-level tasks, focusing on project-level benchmarks that involve multi-class dependencies. A user study with 36 participants examined prompting patterns, screen recordings, and chat logs to identify Human‑LLM Interaction (HLI) features that influence productivity. The results highlight three consistently supportive HLI features, five guidelines to boost productivity, and a taxonomy of 29 runtime and logic errors with mitigation strategies.

By Sangwon Hyun, Hyunjun Kim, Jinhyuk Jang, Hyojin Choi, M. Ali Babar
arXiv AI
Aug 24

Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

The paper investigates how small lexical changes in prompts can cause large performance swings in large language models. Using a dataset of 132,000 prompt variants, the authors uncover a scaling law linking higher average task performance to lower variance and greater robustness. They identify domain-specific terminology and explicit action directives as key linguistic factors that stabilize prompts, and propose an automated Prompt-Refining Agent that reduces performance variance by 40.7% in code generation while maintaining or improving mean performance.

By Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
arXiv AI
6d ago

A Framework for Identifying, Categorizing, and Explaining Bias in AI-Generated Code

The paper presents a taxonomy-driven framework for identifying, categorizing, and explaining bias in AI-generated Python code. By extending an existing dataset and manually annotating bias categories and justifications, the authors evaluate both proprietary and open-source large language models (LLMs) for automated bias detection and explanation. Results show that models such as Gemini and Qwen3-coder achieve high classification accuracy and produce justification and code identification similarities that closely match human-authored reasoning.

By Manaal Basha, Aimee M. Ribeiro, Gema Rodriguez-Perez