arXiv AI By Drew Ross, Arya Hadizadeh Moghaddam, Dongjie Wang, Xiaoyu Zhang, Zijun Yao

SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation

Read the original on arXiv AI →

SoftGene is a new framework that enhances gene set annotation by combining protein language model representations with a hybrid prompting scheme. It uses a hierarchical attention-based encoder built on ESM to encode gene sets from protein amino acid sequences, then merges soft prompts derived from these embeddings with hard prompts generated by a large language model. The approach is evaluated on Gene Ontology and MSigDB datasets, showing that integrating protein-sequence information with textual context improves overall annotation performance, though the benefit varies across biological domains.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 28

Interpreting Latent Protein Language Model Features with Geometric Annotations

The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.

By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong