The paper investigates using synthetic natural-language descriptions to contrastively pretrain small transformer encoders for code representation. By pairing generated descriptions with code in a dual-encoder setup during training and discarding them at inference, the authors achieve significant improvements over traditional pretraining baselines on most evaluated tasks. When fine‑tuned, these models match or surpass much larger zero‑shot models and remain competitive with execution‑aware supervision, indicating a scalable alternative for code embeddings.
By Kenneth Paulsen, Florian Tambon, Mike Papadakis, Shin Yoo
arXiv:2508. 21290v2 Announce Type: replace-cross Abstract: jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical question-answering, and identify semantically similar code snippets across programming languages.
By Daria Kryvosheieva, Saba Sturua, Michael G\"unther, Han Xiao
arXiv:2608. 15412v1 Announce Type: cross Abstract: Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive.
By Yifeng He, Yundi Xu, Christopher Castro Gaw Gonzalo, Zili Wang, Hao Chen
The paper introduces LWVIC4Code, a non‑contrastive, layer‑wise representation learning method for detecting Type‑IV code clones—semantically equivalent fragments that differ syntactically. It builds on the VICReg framework, adding cross‑layer consistency regularization and depth‑dependent weighting to refine semantic information across transformer layers. Experiments on Python and multi‑language datasets show that LWVIC4Code matches or outperforms contrastive baselines and zero‑shot large language models, generalizing well to Java and C# without requiring negative samples.
By Luciano Marchezan, Kevin Delcourt, Eugene Syriani, Houari Sahraoui
The paper investigates whether training small decoder-only transformers on code‑switched text can induce cross‑lingual alignment. Using two 100‑million‑word multilingual corpora—one a mix of English, Dutch, and Chinese BabyBabelLM data, and another generated by inserting word‑ and sentence‑level code‑switching via an LLM—the authors find that code‑switched training aligns representations of parallel text, especially across different scripts, and that this alignment persists when later training on monolingual documents. A curriculum that progresses from word‑level code‑switching to sentence‑level code‑switching and finally to monolingual data yields models that outperform baselines on the BabyLM evaluation suite, demonstrating that code‑switching curriculum learning is an effective data augmentation strategy for multilingual pretraining.
By Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley
The paper introduces LSem2Vec, a two‑stage method that first uses a large language model to extract source code semantics and then applies a sentence embedding model to produce vector representations. This approach removes the need for task‑specific training or fine‑tuning, addressing errors in LLM outputs. Experiments on three datasets across multiple programming languages show that LSem2Vec outperforms five state‑of‑the‑art unsupervised methods.
By Zixiang Xian, Chenhui Cui, Rubing Huang, Chunrong Fang, Zhenyu Chen
arXiv:2608. 07894v1 Announce Type: new Abstract: Recent advances have highlighted the potential of machine learning, particularly Large Language Models (LLMs), for analyzing and optimizing programs.
By Calvin Higgins, Marco Alvarez
The paper introduces geometric iterative retrieval, a new approach for resynthesizing high‑quality audio from coarse Residual Vector Quantization (RVQ) codec tokens. Instead of choosing between discrete token prediction or continuous regression, the method performs contrastive retrieval within the continuous codebook space, leveraging the RVQ hierarchy as an iterative decomposition. Experiments on speech and music codec restoration tasks demonstrate that this technique outperforms both single‑pass token prediction and one‑step regression baselines.
By Leo Schmidt-Traub, Fr\'ed\'eric Berdoz, Luca A. Lanzend\"orfer, Roger Wattenhofer
LLM-based text embedders have substantially improved retrieval and semantic representation quality, but their deployment remains costly: large backbone models slow down embedding inference, while high-dimensional full-precision embeddings impose substantial storage and bandwidth overhead on large-scale indexes. In this paper, we present BITEMBED, an extreme low-bit framework for LLM-based text embedding that jointly targets encoding efficiency and vector storage.