arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
By Sarwan Ali
arXiv:2609.14882v1 Announce Type: cross
Abstract: Nucleotide sequence analysis is central to problems spanning regulatory genomics, evolutionary biology, and phenotype prediction. Classical bioinform...
By Evgeny S. Saveliev, Krzysztof Kacprzyk, Charlotte Capitanchik, Neelanjan Mukherjee, Kate Matlin, Ryan Sheridan, Srinivas Ramachandran, Jernej Ule, David L. Bentley, Mihaela van der Schaar
arXiv:2607. 00931v1 Announce Type: new Abstract: Predicting cancer drug response from transcriptomic profiles is a cornerstone of precision oncology, yet the scientific value of machine learning models hinges not solely on predictive accuracy, but also on their capacity to generate reliable biological insights.
By Martino Ciaperoni, Margherita Lalli, Simone Piaggesi, Martina Varisco, Francesco Carli, Riccardo Guidotti, Dino Pedreschi, Francesco Raimondi, Fosca Giannotti
The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.
By Michael Bouzinier, Dmitry Etin
arXiv:2512. 22240v5 Announce Type: replace-cross Abstract: Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biological findings.
By Chama Bensmail
arXiv:2606. 01042v1 Announce Type: cross Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions.
By Xinyu Yuan, Xixian Liu, Jianan Zhao, Yashi Zhang, Hongyu Guo, Jian Tang
arXiv:2601. 12805v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks.
By Xiaohan Huang, Meng Xiao, Chuan Qin, Qingqing Long, Jinmiao Chen, Yuanchun Zhou, Hengshu Zhu
arXiv:2608. 05359v1 Announce Type: new Abstract: CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, exposed via MCP.
By Jose A. Bird
The paper introduces OmicsBench, a new reasoning benchmark for multi‑omics sequences that includes 1,160 expert‑validated questions across DNA regulation, RNA processing, and protein function tasks, requiring traceable evidence chains. Evaluation of 17 large language models shows that scientific LLMs, while more accurate in classification, often lack valid evidence, suggesting shortcut learning. To address this, the authors propose tool‑augmented on‑policy distillation (TA‑OPD), a post‑training method that improves both evidence grounding and predictive performance across five Qwen3.5 models of varying sizes.
By Jie Ying, Zhefan Wang, Zihong Chen, Zhengqing Li, Jinzhe Li, Gang Li, Jian Liu, Fang Hu, Tao Luo, Zhonghang Yuan, Wanli Ouyang, Stan Z. Li, Fan Yang, Nanqing Dong
OpenMTB‑Audit is an open‑source benchmark that tests large language models on 500 synthetic non‑small cell lung cancer cases, covering five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. The study found that all eight tested LLMs over‑refused Partially Supported recommendations, collapsing labels to achieve high safety scores. A deterministic seven‑module framework, MTB‑AuditAgent, was introduced to reduce over‑refusal to 6.7% and reach 91.2% accuracy, while an oncologist annotation study highlighted disagreement around the boundary between information sufficiency and treatment optimization.
By Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou
arXiv:2509.20702v3 Announce Type: replace-cross
Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
By Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou, Zhun Deng, Haoyu Zhang, Xihao Li, Didong Li
arXiv:2606. 06224v1 Announce Type: cross Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology.
By Yanqing Luo (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Julius Hense (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany), Niklas Preni{\ss}l (Institute of Pathology, Charit\'e Universit\"atsmedizin, Berlin, Germany, Berlin Institute of Health at Charit\'e -- Universit\"atsmedizin Berlin, BIH Biomedical Innovation Academy, BIH Charit\'e Digital Clinician Scientist Program, Berlin, Germany), Andreas Mock (Institute of Pathology, Ludwig Maximilian University of Munich, Munich, Germany, Division of Translational Medical Oncology, DKFZ, Heidelberg, Germany, NCT Heidelberg, Heidelberg, Germany, German Cancer Consortium), Klaus-Robert M\"uller (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany, Department of Artificial Intelligence, Korea University, Seoul, Korea, Max-Planck Institute for Informatics, Saarbr\"ucken, Germany), Thomas Schnake (Department of Chemistry, Chemical Physics Theory Group, University of Toronto, Canada, Vector Institute for Artificial Intelligence, Toronto, Canada, Acceleration Consortium, University of Toronto, Canada), Mina Jamshidi Idaji (Berlin Institute for the Foundations of Learning and Data, Berlin, Germany, Machine Learning Group, Technische Universit\"at Berlin, Berlin, Germany)