arXiv AI By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Read the original on arXiv AI →

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Tool-Augmented On-Policy Distillation for LLM Domain Adaptation in Sequence-Based Omics Tasks

The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.

By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
arXiv AI
Sep 15

Towards a knowledge-enhanced single-cell foundation model

The paper introduces scKITE, a single-cell foundation model that incorporates biological knowledge—cell-level text annotations and gene-level regulatory information—into a shared Transformer encoder via lightweight auxiliary decoders used only during pretraining. This approach provides a new scaling dimension beyond merely increasing data size, enabling the model to achieve superior performance on diverse downstream tasks with only 179,067 pretraining samples, less than 0.5% of the data used by previous strong scFMs. The study demonstrates that knowledge-enhanced pretraining can yield significant gains while reducing computational cost.

By Hanqing Zhang, Jie Bao, Mei Ma, Shuai Liu, Jiaying Ma, Jiaguan Liu, Jiaxiao Li, Zhenbo Li, Wenwen Gong, Zhijun Ca