arXiv AI

Invariant Pretraining for Robust Code Representations

arXiv:2608. 15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive.

arXiv Machine Learning
Sep 16

Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

The paper introduces LWVIC4Code, a non‑contrastive, layer‑wise representation learning method for detecting Type‑IV code clones—semantically equivalent fragments that differ syntactically. It builds on the VICReg framework, adding cross‑layer consistency regularization and depth‑dependent weighting to refine semantic information across transformer layers. Experiments on Python and multi‑language datasets show that LWVIC4Code matches or outperforms contrastive baselines and zero‑shot large language models, generalizing well to Java and C# without requiring negative samples.

By Luciano Marchezan, Kevin Delcourt, Eugene Syriani, Houari Sahraoui
arXiv AI
Sep 4

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

The paper investigates using synthetic natural-language descriptions to contrastively pretrain small transformer encoders for code representation. By pairing generated descriptions with code in a dual-encoder setup during training and discarding them at inference, the authors achieve significant improvements over traditional pretraining baselines on most evaluated tasks. When fine‑tuned, these models match or surpass much larger zero‑shot models and remain competitive with execution‑aware supervision, indicating a scalable alternative for code embeddings.

By Kenneth Paulsen, Florian Tambon, Mike Papadakis, Shin Yoo
arXiv AI
2d ago

Are AI Coders Snitches? An Empirical Study of Pretraining Data Detection on Code Large Language Models

The paper investigates whether code large language models (CodeLLMs) inadvertently reproduce proprietary or sensitive code by evaluating seven state‑of‑the‑art training data detection (TDD) methods on eight CodeLLMs. It introduces CodeSnitch, a benchmark of 9,000 function‑level code samples across three languages, each labeled as included or excluded from training data, and applies mutation strategies based on the Type‑1 to Type‑4 code clone taxonomy to test TDD robustness. The study offers a systematic assessment of current TDD techniques for code and suggests directions for developing more effective detection methods.

By Tianlin Li, Yunxiang Wei, Zhiming Li, Aishan Liu, Qing Guo, Xianglong Liu, Dongning Sun, Yang Liu
arXiv AI
Jun 24

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.

By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv Computation and Language
6d ago

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

The paper investigates whether training small decoder-only transformers on code‑switched text can induce cross‑lingual alignment. Using two 100‑million‑word multilingual corpora—one a mix of English, Dutch, and Chinese BabyBabelLM data, and another generated by inserting word‑ and sentence‑level code‑switching via an LLM—the authors find that code‑switched training aligns representations of parallel text, especially across different scripts, and that this alignment persists when later training on monolingual documents. A curriculum that progresses from word‑level code‑switching to sentence‑level code‑switching and finally to monolingual data yields models that outperform baselines on the BabyLM evaluation suite, demonstrating that code‑switching curriculum learning is an effective data augmentation strategy for multilingual pretraining.

By Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley
arXiv Machine Learning
Sep 10

LLM Layers Immediately Correct Each Other

arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...

By Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt