On the Robustness of LLMs' Internal Representation of Code Correctness
arXiv:2608. 08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally.
arXiv:2606. 28574v1 Announce Type: cross Abstract: When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder.
arXiv:2608. 08266v1 Announce Type: cross Abstract: Code generated by modern language models often reads naturally.
arXiv:2607. 08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value.
arXiv:2607. 08731v2 Announce Type: replace-cross Abstract: National language models are becoming publicly funded epistemic infrastructure.
arXiv:2609.14754v1 Announce Type: cross Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
arXiv:2604. 03447v2 Announce Type: replace-cross Abstract: LLM-based software engineering assistants often reason over multiple artifacts, including code, documentation, signatures, and tests, even when those artifacts are incomplete or mutually inconsistent.
Large language models (LLMs) are increasingly used as judges for code evaluation, assessing correctness without reference implementations. This study investigates whether LLM judges can fairly evaluate semantically equivalent code that differs in superficial aspects such as variable names, comments, or formatting. The authors define six types of potential bias, conduct experiments across five programming languages and multiple LLMs, and find that all tested judges exhibit both positive and negative biases, leading to inflated or unfairly low scores even when prompted to generate test cases.
arXiv:2608. 19009v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly paired with verifiers (step checkers, self-consistency filters, tool-based fact checkers, formal proof assistants) that claim to detect the model's errors.
arXiv:2609.21190v1 Announce Type: cross Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check corr...
The paper "SoK: Formal Methods for Fact-Checking and Information Integrity" discusses how automated fact‑checking systems typically output a verdict but lack a detailed record—called a warrant—explaining the evidence and conditions behind that verdict. It proposes organizing the field by what is being formalised—claims, reasoning, checking systems, ecosystems, and regulatory obligations—rather than by pipeline stages, and surveys 121 works to identify gaps, notably the scarcity of formal methods applied to verifying the checking systems themselves. The authors highlight that existing formal tools, though largely unused in this domain, could address these gaps and outline open problems with suggested first steps.
arXiv:2601.08654v3 Announce Type: replace Abstract: Rubric-based text evaluation increasingly relies on large language models (LLMs) as scalable judges, yet frozen black-box models can interpret the...
arXiv:2610.01847v1 Announce Type: cross Abstract: Model specifications define how large language models (LLMs) should behave, guiding alignment training, inference-time behavior, and evaluation. Yet...
arXiv:2510. 07315v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check.