ProteinJEPA introduces a joint‑embedding predictive architecture that supplements masked language modeling (MLM) with a cosine loss to predict latent representations of a teacher model. On 19 protein tasks, MLM+JEPA outperforms compute‑matched and step‑matched MLM‑only training across 78 and 76 of 114 comparisons, achieving notable gains on structure‑ and homology‑sensitive tasks such as SCOPe‑40 retrieval and remote homology. Ablation studies show the cosine loss is superior to mean squared error and that latent prediction complements rather than replaces MLM.
By Dan Ofer, Dafna Shahaf, Michal Linial
arXiv:2608.20647v1 Announce Type: new
Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (...
By Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
arXiv:2608. 03629v1 Announce Type: new Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried additively through a residual stream.
By Abdallah Khemais
arXiv:2606. 00091v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) have reshaped self-supervised representation learning in vision.
By Sangdae Nam
arXiv:2508. 08289v3 Announce Type: replace Abstract: Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave.
By Mu Qiao
arXiv:2606. 05169v1 Announce Type: new Abstract: We give a stereological theory of LLM benchmark coverage.
By Jason Z Wang