The paper presents MoCoP v2, an enhanced contrastive pretraining method that aligns small molecule embeddings with deep‑learning‑derived cell morphology profiles. By replacing CellProfiler fingerprints with richer image‑encoded features, the new embeddings better capture how molecules alter cell morphology, leading to improved QSAR, toxicity, ADME, and activity predictions. Performance scales log‑linearly with training data size, indicating further gains with larger datasets.
By Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi, Zhizhuo Zhang
The paper proposes a three‑stage training pipeline that begins with procedural pretraining on abstract, procedurally generated data, followed by molecular pretraining on SMILES, and finally downstream fine‑tuning for molecular property prediction. Experiments show that procedural pretraining improves downstream performance—e.g., a 4.8% error reduction on Lipophilicity—especially when labeled data are scarce, and that the benefit peaks at an intermediate procedural training budget. Analysis indicates that transferable knowledge resides mainly in attention layers, while feed‑forward layers may over‑specialize.
By Moritz Friedemann, Zachary Shinnick, Philip Torr, Bruno Andreis
arXiv:2507. 04704v3 Announce Type: replace-cross Abstract: Understanding how cellular morphology, gene expression, and spatial context jointly shape tissue function is a central challenge in biology.
By Zhenglun Kong, Mufan Qiu, John Boesen, Xiang Lin, Sukwon Yun, Tianlong Chen, Manolis Kellis, Marinka Zitnik
arXiv:2606. 01461v1 Announce Type: new Abstract: Developing effective anticancer therapeutics remains challenging due to tumor heterogeneity and the absence of well-defined molecular targets across cancer subtypes.
By Brenda Nogueira, Gisela A. Gonzalez-Montiel, Nitesh V. Chawla, Nuno Moniz
arXiv:2511. 19264v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting adoption in drug discovery, where chemists need interpretable rationales for proposed structures.
By Amirtha Varshini A S, Duminda S. Ranasinghe, Hok Hei Tam
arXiv:2606. 08802v1 Announce Type: new Abstract: Standard flow and diffusion pre-training matches the distribution of available data (e.
By Riccardo De Santi, Bruce Lee, Cristian Perez Jensen, Kimon Protopapas, Sophia Tang, Cheng-Hao Liu, Pranam Chatterjee, Yisong Yue, Andreas Krause
arXiv:2608.07632v2 Announce Type: replace-cross
Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Pa...
By Al\'an F. Mu\~noz, Johan Fredin Haslum, Runxi Shen, Anne E. Carpenter, Shantanu Singh
The paper introduces Distribution‑Conditioned Transport (DCT), a framework that learns transport maps conditioned on embeddings of source and target distributions, allowing generalization to unseen distribution pairs. DCT supports semi‑supervised learning for distributional forecasting by leveraging distributions observed at only one condition. It is agnostic to the transport mechanism and is demonstrated on synthetic benchmarks and four biological applications, including batch effect transfer in single‑cell genomics and modeling T‑cell receptor sequence evolution.
By Nic Fishman, Gokul Gowri, Paolo L. B. Fischer, Marinka Zitnik, Omar Abudayyeh, Jonathan Gootenberg
arXiv:2606. 11508v1 Announce Type: new Abstract: Accurate prediction of absorption, distribution, metabolism, and excretion (ADME) properties is critical to drug discovery, but remains challenging because ADME endpoints are noisy, interdependent, and often data-limited.
By Yifan Xue, Srimukh Prasad Veccham, Saee Paliwal, Tyler Shimko, Micha Livne
arXiv:2603. 13377v2 Announce Type: replace-cross Abstract: Representation learning has driven major advances in natural image analysis by enabling models to acquire high-level semantic features.
By Ivan Svatko, Maxime Sanchez, Ihab Bendidi, Gilles Cottrell, Auguste Genovesio
arXiv:2606. 31126v1 Announce Type: new Abstract: Predicting biomolecular properties from limited labeled data is a central bottleneck in protein engineering and small-molecule design.
By Davy Guan, Lu Zhang, Asiri Wijesinghe, Allen Zhu, He Zhao, Helen Power, F. Hafna Ahmed, Andrew Warden, Cheng Soon Ong, Daniel M. Steinberg
SCALE is a conditional transport model that treats cells as unordered sets to predict treated cell populations without requiring cell-level matching. It uses a shared set-aware encoder and a conditional DiT backbone to learn latent transport, enabling endpoint supervision that is directly delta-aligned. Across diverse perturbation types—including genetic, chemical, developmental, and immune—SCALE accurately recovers gene‑expression changes, response directions, and population structure, outperforming competing methods on CRISPR data and successfully prioritizing cytokines that elicit distinct immune responses.
By Shuizhou Chen, Lang Yu, Xueqin Lin, Xinjie Mao, Songming Zhang, Xinyu Gu, Hao Wu, Sheng Xu, Kedu Jin, Lei Bai, Quan Qian, Qin Chen, Qiang Gao, Siqi Sun, Zhangyang Gao