arXiv Machine Learning By Sarwan Ali, Taslim Murad, Imdadullah Khan, Safi Faizullah

When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification

Read the original on arXiv Machine Learning →

The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

ProteinJEPA: Latent prediction improves protein language model pretraining

ProteinJEPA introduces a joint‑embedding predictive architecture that supplements masked language modeling (MLM) with a cosine loss to predict latent representations of a teacher model. On 19 protein tasks, MLM+JEPA outperforms compute‑matched and step‑matched MLM‑only training across 78 and 76 of 114 comparisons, achieving notable gains on structure‑ and homology‑sensitive tasks such as SCOPe‑40 retrieval and remote homology. Ablation studies show the cosine loss is superior to mean squared error and that latent prediction complements rather than replaces MLM.

By Dan Ofer, Dafna Shahaf, Michal Linial
arXiv AI
Jul 23

Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models

arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.

By Sarwan Ali