arXiv AI By Manuel Pita

Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs

Read the original on arXiv AI →

arXiv:2606. 28574v1 Announce Type: cross Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 24

Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

Large language models (LLMs) are increasingly used as judges for code evaluation, assessing correctness without reference implementations. This study investigates whether LLM judges can fairly evaluate semantically equivalent code that differs in superficial aspects such as variable names, comments, or formatting. The authors define six types of potential bias, conduct experiments across five programming languages and multiple LLMs, and find that all tested judges exhibit both positive and negative biases, leading to inflated or unfairly low scores even when prompted to generate test cases.

By Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang, Yongil Kim, Kyomin Jung