arXiv:2606. 26563v1 Announce Type: cross Abstract: Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integration of metadata, assay context, and auxiliary evidence.
By Ian Diks, Zhen Yang, Arjun Banerjee, Tim Proctor, Kenny Workman
The study evaluated nine transcriptomic models—five bulk RNA‑seq and four single‑cell RNA‑seq—designed to predict response to immune checkpoint inhibitors. Across independent datasets, bulk models performed near chance while single‑cell models offered only modest gains, and pathway analyses revealed inconsistent biomarker signals. The results highlight the limited cross‑cohort robustness and biological consistency of current transcriptomic ICI predictors.
By Yuheng Liang, Lucy Chhuo, Ahmadreza Argha, Nona Farbehi, Lu Chen, Roohallah Alizadehsani, Mehdi Hosseinzadeh, Min Yang, Thantrira Porntaveetusm, Youqiong Ye, Hamid Alinejad-Rokny
arXiv:2603. 11872v3 Announce Type: replace-cross Abstract: Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI systems lack direct access to transcriptomic representations while expression foundation models remain opaque to natural language.
By Omar Coser
arXiv:2606. 29949v1 Announce Type: cross Abstract: H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability.
By Dominik Winter, Dominik Vonficht, Lo\"ic Le Bescond, Christian Gebbe, Marco Rosati, Richard J. Chen, Markus Schick, Ross Stewart, Nicolas Brieu
arXiv:2602. 17330v5 Announce Type: replace-cross Abstract: Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost of pairwise affinity evaluations and dataset imbalances that obscure clinically important minority clonotypes.
By Rong Fu, Zijian Zhang, Kun Liu, Jiekai Wu, Xianda Li, Simon Fong
arXiv:2509.20702v3 Announce Type: replace-cross
Abstract: Recent advances in large language model (LLM) embeddings have enabled powerful representations for biological data, but most applications to...
By Hongqian Niu, Jordan Bryan, Jacob Williams, Hufeng Zhou, Zhun Deng, Haoyu Zhang, Xihao Li, Didong Li
arXiv:2609.14709v1 Announce Type: new
Abstract: Immune therapies act across cell-intrinsic programs, tissue ecosystems, and patient-specific immune states, yet most predictors address these scales se...
By Taoyong Cui, Xi Wang, Zonghang Li, Jinchao Ding, Lingsen You, Yuzhi Xu, Wanghan Xu, Fang Wu, Kejun Ying, Wanli Ouyang, Pheng Ann Heng, Ling Yang, Zhenfei Yin, Yingcheng Wu
arXiv:2608.24688v1 Announce Type: new
Abstract: Precision oncology necessitates a longitudinal model of patient state that captures cancer evolution and treatment over time, integrating multimodal ob...
By Eugene Vorontsov, Yi Kan Wang, Alican Bozkurt, Adam Casson, Ludmila Tydlitatova, Michal Zelechowski, Ezra E. W. Cohen, Jyoti D. Patel, Max Banaszak, Caitlin McWilliams, Shane Colley, Kate Sasser, Ryan Fukushima, Eric Lefkofsky, Razik Yousfi, Siqi Liu
arXiv:2602. 01051v5 Announce Type: replace Abstract: Repertoire-level analysis of T cell receptors offers a biologically grounded signal for disease detection and immune monitoring, yet practical deployment is impeded by label sparsity, cohort heterogeneity, and the computational burden of adapting large encoders to new tasks.
By Rong Fu, Muge Qi, Yang Li, Yabin Jin, Jiekai Wu, Chunlei Meng, Juntao Gao, Li Bao, Qi Zhao, Wei Luo, Youjin Wang, Simon Fong
The paper presents a newly curated, multi-center, multi-modal, and longitudinal lung cancer dataset comprising 1,365 patients with whole-slide images, CT scans, PET scans, structured clinical data, transcriptomics, and follow-up information. The dataset features substantial, non-uniform missingness across modalities, making it ideal for evaluating robust multi-modal fusion strategies. Benchmarks on 12‑month overall survival, disease‑specific survival, and longitudinal hazard prediction demonstrate that integrating complementary modalities consistently outperforms uni-modal approaches, even under severe missing data.
By Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago, Catarina Barata
NeoRed is a multimodal large language model specifically designed for diagnosing neonatal respiratory diseases. It addresses two main limitations of existing models—domain gaps from adult data and inadequate integration of clinical context—by leveraging two real-world neonatal datasets (NeoCXR and NeoCXR-EV). The model incorporates a Knowledge-Logic-Alignment framework that injects diagnostic priors, aligns report semantics with diagnostic logic, and aligns visual features with imaging conclusions, achieving superior performance on neonatal benchmarks while maintaining adult benchmark performance.
By Yinan Liu, Hongtai Xia, Haoran Xu, Jiankang Hong, Jingkuan Song, Ye Luo
arXiv:2511. 03354v2 Announce Type: replace-cross Abstract: Generative artificial intelligence (GenAI) is transforming bioinformatics by advancing genomics, proteomics, transcriptomics, structural biology, and drug discovery.
By Wasimul Karim, Riasad Alvi, Sayeem Been Zaman, Arefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Md Rafi Ur Rashid, Md Rafiqul Islam, Yakub Sebastian, Sami Azam