arXiv Computation and Language

QLoRA Fine-Tuning of Ministral LLM for Sequence-to-Function Protein Annotation

arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv Machine Learning
Sep 17

Q-BIOLAT: Binary Latent Protein Fitness Landscapes for QUBO-Based Optimization

Q-BIOLAT is a framework that converts pretrained protein-language-model embeddings into compact binary codes and trains a quadratic unconstrained binary optimization (QUBO) surrogate with unary and pairwise latent interactions for protein fitness optimization. The study demonstrates that binary encodings with similar predictive accuracy can produce different Hamming neighborhoods, affecting local optima and search trajectories, and shows that PCA followed by per‑coordinate median thresholding yields a more balanced binary space than AE/VAE baselines. Experimental evaluation on GFP and AAV fitness landscapes from ProteinGym confirms that simulated annealing, genetic algorithms, and greedy hill climbing can retrieve high‑percentile variants, with decoded candidates reported via surrogate‑predicted scores.

By Truong-Son Hy
arXiv AI
Aug 24

Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

The study explores using large language models (LLMs) to create ranking policies for shortlisting protein binders after they have been generated by de novo design workflows. By combining precomputed structural‑confidence and interface‑quality proxy scores, the authors demonstrate that iterative LLM policies can modestly improve recall and NDCG metrics over single‑feature baselines. The approach offers an interpretable post‑generation decision layer that helps prioritize binders from large candidate pools.

By Gyubok Lee, Kiwoong Yoo, Jimin Seo, Kyunghoon Hur, Edward Choi
arXiv AI
Sep 25

PFArena: Benchmarking Language Models for Protein Modification

PFArena is a new benchmark for evaluating language models in protein modification tasks, featuring four controlled interfaces that span single‑mutant generation and multi‑mutant ranking. It incorporates varying levels of mutation fitness data to represent four research scenarios with different amounts of prior experimental context. The benchmark tests six protein language models, six large language models, and five LLM‑based agents, finding that PLMs excel at open‑ended single‑mutant generation while LLMs and agents perform best in multi‑mutant ranking when target‑specific data are available, yet all struggle as search space and mutation depth grow.

By Yawen Ouyang, Xinbo Zhang, Ziyuan Ma, Yixin Wu, Wenbin Liao, Feiran Zhang, Wenjie Li, Lihao Wang, Hao Wang, Xiaoqing Zheng, Xuefeng Yan, Lei Bai, Ya-Qin Zhang, Shuyi Zhang, Wei-Ying Ma, Dahua Lin, Bowen Zhou, Hao Zhou