arXiv AI

Protein contacts are already in the attention: a single-forward-pass alternative to the Categorical Jacobian

arXiv:2606. 21876v2 Announce Type: replace-cross Abstract: The Categorical Jacobian of Zhang et al.

arXiv AI
Sep 24

ProteinJEPA: Latent prediction improves protein language model pretraining

ProteinJEPA introduces a joint‑embedding predictive architecture that supplements masked language modeling (MLM) with a cosine loss to predict latent representations of a teacher model. On 19 protein tasks, MLM+JEPA outperforms compute‑matched and step‑matched MLM‑only training across 78 and 76 of 114 comparisons, achieving notable gains on structure‑ and homology‑sensitive tasks such as SCOPe‑40 retrieval and remote homology. Ablation studies show the cosine loss is superior to mean squared error and that latent prediction complements rather than replaces MLM.

By Dan Ofer, Dafna Shahaf, Michal Linial
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
arXiv Machine Learning
Aug 31

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

The paper introduces DARTS, a method for tuning decoder representations during model merging. It addresses representation bias in autoregressive decoders by using an entropy‑weighted L1 loss and a per‑position additive bias to correct errors that accumulate across token positions. Experiments on code generation, mathematical reasoning, and instruction following with Llama‑2‑7B show that DARTS improves performance over standard surgery while adding only 0.1% extra parameters.

By Aaryan Ajay Sharma, Sai Nishanth Padala, Seganrasan Subramanian
arXiv Computation and Language
3d ago

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov