arXiv:2606. 01042v1 Announce Type: cross Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions.
By Xinyu Yuan, Xixian Liu, Jianan Zhao, Yashi Zhang, Hongyu Guo, Jian Tang
arXiv:2608. 16419v1 Announce Type: cross Abstract: Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces.
By Zhenchao Tang, Xiaogang Xu, Tianxu Lv, Jiahui Guan, Jiale Zhou, Haohuai He, Zhi Song, Hanbo Huang, Jiehui Huang, Jiafei Wu, Zhe Liu
arXiv:2607. 18777v1 Announce Type: new Abstract: Evaluating machine learning in scientific domains requires separating correct predictions from correct reasons under realistic distribution shifts.
By Dongkwan Kim, Yiming Gao, Yining Yang, Yang Shen
arXiv:2608. 12906v1 Announce Type: cross Abstract: RNA-Protein Interactions (RPIs) are critical for regulating cellular functions.
By Danyu Li, Ling Zhou, Rubing Huang, Xian Zhong, Bin Zou, Kui Jiang
arXiv:2606. 27752v1 Announce Type: new Abstract: Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions.
By Dongxia Wu, Mingyu Li, Yuhui Zhang, Anurendra Kumar, Emma Lundberg, Serena Yeung-Levy, Emily B. Fox
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis
The paper introduces LitEm, a neural regression model that allows transductive knowledge graph embedding models to predict numerical attributes. LitEm achieves top or near‑top performance on most attributes across datasets such as FB15K‑237, YAGO15K, DB15K, and Mutagenesis. A co‑training framework further improves link prediction for bilinear models while enabling them to predict numerical attributes, demonstrating literal‑aware encoding of attribute information.
By Rupesh Sapkota, Louis Mozart Kamdem Teyou, Moshood Yekini, Caglar Demir, Axel-Cyrille Ngonga Ngomo
arXiv:2606. 08147v1 Announce Type: cross Abstract: DNA cis-regulatory elements (CREs) such as enhancers control gene expression levels.
By Yi Duan, Zhao Yang, Jiwei Zhu, Ying Ba, Chuan Cao, Bing Su
arXiv:2608. 05928v1 Announce Type: new Abstract: Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes.
By Yuhao Wang, Zelin Zang, Yuxuan Liu, Zhen Lei, Stan Z. Li
arXiv:2603. 10395v2 Announce Type: replace Abstract: Graph generation is a fundamental task with broad applications, such as drug discovery.
By Baoheng Zhu, Deyu Bo, Delvin Ce Zhang, Xiao Wang
arXiv:2602. 07832v3 Announce Type: replace-cross Abstract: Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward hacking.
By Xian Wu, Kaijie Zhu, Ying Zhang, Lun Wang, Wenbo Guo
arXiv:2609.06779v1 Announce Type: cross
Abstract: Drug repurposing aims to identify new therapeutic uses for existing compounds and, compared with de novo drug discovery, offers a faster and more cos...
By Zijie Liu, Hongxuan Li, Zhen Tan, Jinhao Duan, Baixiang Huang, Zunpeng Liu, Kai Shu, Tianlong Chen