CurvFlow-DTA introduces a dual-graph discrete Ricci curvature flow framework for drug–target affinity prediction, replacing static curvature with weighted Forman curvature flow on both drug and protein residue–residue contact graphs. The method precomputes a label‑independent flow trajectory for each entity and uses a pair‑conditioned selector to guide a dual‑branch Flow‑GINE, leveraging frozen ESM‑2 residue representations. Experiments on Davis and KIBA datasets show significant improvements over the Ricci‑GraphDTA baseline, with reductions in mean squared error of up to 19.9% in warm‑start and 27.4% in cold‑start settings, and higher concordance indices across benchmarks.
By Jicheng Ma, Yunyan Yang, Juan Zhao, Liang Zhao
The paper introduces a scalable method to interpret sparse autoencoder (SAE) features in the ESM-2 protein language model by leveraging geometrically inspired features of the protein α‑carbon backbone. Across 8M layers of ESM-2, a false discovery rate–controlled analysis shows that local geometry is significantly associated with many SAE features, revealing substructure within known biological labels and enabling annotation of unannotated metagenomic proteins. Ablation experiments demonstrate that removing these geometric features shifts ESM-2’s predicted contact maps toward the descriptor, linking mechanistic interpretability with structural biology.
By Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
arXiv:2608.29207v1 Announce Type: new
Abstract: Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (t...
By Yifan Feng, Guanjie Cheng, Shihui Ying, Shaoyi Du, Yue Gao
arXiv:2605. 01625v3 Announce Type: replace Abstract: Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial resolutions, ranging from atomic interactions to global fold topology.
By Viet Thanh Duy Nguyen, John K. Johnstone, Truong-Son Hy
arXiv:2607. 22777v1 Announce Type: cross Abstract: Protein language models learn transferable sequence representations.
By Chen Wang, Boming Kang, Qinghua Cui
The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.
By Sarwan Ali, Taslim Murad, Imdadullah Khan, Safi Faizullah