Task- and dataset-specific information in protein language models
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.
arXiv:2608. 12090v1 Announce Type: new Abstract: Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology.
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs).
arXiv:2606. 27242v1 Announce Type: new Abstract: Training-free source selection for LLM families with shared vocabularies arises in scientific string domains such as SMILES, protein, and genomic sequences, where candidate corpora share a tokenizer but differ in prediction targets.
arXiv:2603. 14717v2 Announce Type: replace Abstract: Generating novel protein sequences that respect a family's statistical constraints typically requires training deep generative models on thousands to millions of examples.
ProteinJEPA introduces a joint‑embedding predictive architecture that supplements masked language modeling (MLM) with a cosine loss to predict latent representations of a teacher model. On 19 protein tasks, MLM+JEPA outperforms compute‑matched and step‑matched MLM‑only training across 78 and 76 of 114 comparisons, achieving notable gains on structure‑ and homology‑sensitive tasks such as SCOPe‑40 retrieval and remote homology. Ablation studies show the cosine loss is superior to mean squared error and that latent prediction complements rather than replaces MLM.
arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
The paper investigates a problem in guided protein language models where strong guidance causes the model’s internal representations to collapse onto a region indistinguishable from random amino‑acid input, leading to low‑complexity sequences that still score well on the targeted property. The authors identify this off‑manifold collapse as a detectable signature and propose a post‑hoc filtering technique—Mahalanobis filtering—that removes atypical candidates based on a density prior over natural activations. This simple, training‑free step improves both property scores and structural plausibility across different guidance methods without altering the generator.
The paper investigates how guided protein language models can collapse onto off‑manifold representations when heavily steered to optimize a property. This collapse causes generated sequences to become low‑complexity and statistically similar to random amino‑acid input, yet the property oracle may still rate them highly. The authors propose a cheap, training‑free Mahalanobis filtering step that removes such off‑manifold candidates, improving both property scores and structural plausibility without altering the generator.
arXiv:2605. 16331v2 Announce Type: replace-cross Abstract: Protein language models are increasingly used to guide experimental and clinical decisions, yet it is often unclear whether a confident prediction reflects recognition of biological evidence or retrieval of a statistical default.
arXiv:2609.37675v1 Announce Type: new Abstract: Protein Language Models (PLMs) have made remarkable progress following scaling laws established in natural language processing across sequence- and str...
The paper introduces Speculative Probing, a method that repurposes the speculative‑decoding module of large language models for real‑time classification tasks. By appending a trained soft prompt to the target sequence, the approach leverages the already‑cached KV store during inference, adding negligible overhead while achieving higher accuracy than traditional hidden‑state probes. Experiments on four classification tasks across multiple models show that these lightweight probes outperform zero‑shot GPT‑5.4‑mini and rival or surpass specialized 8B safety classifiers without running a full LLM.
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.