GyroNovo is a new framework for de novo peptide sequencing that improves fragment imputation by guiding the process with decoder errors observed during training. It introduces mass-aware attention using rotary embeddings to encode pairwise mass differences between spectral peaks, and creates easy and hard augmented views of spectra to train the decoder under varying corruption levels. Experiments on NovoBench demonstrate significant gains, with about 9 percentage points higher peptide-level precision and 7 percentage points higher amino-acid-level precision compared to the state-of-the-art baseline.
By Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed
arXiv:2602. 22822v3 Announce Type: replace Abstract: Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice.
By Yunhua Zhong, Yixuan Tang, Yifan Li, Pan Liu, Zhiwen Yang, Jie Yang, Jun Xia
arXiv:2607. 23607v1 Announce Type: new Abstract: Molecular structure elucidation from tandem mass spectra (MS/MS) is a central inverse problem in analytical chemistry.
By Xin Zhao, Yumin Liu, Zhuo Li, Weichu Zheng, Feng Zhu, Xiaokang Yang, Yaohui Jin, Yanyan Xu
arXiv:2608.30175v1 Announce Type: new
Abstract: Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed target...
By Jiaxin Tian, Darren An, Jun Li
MT-ProtBERT is a multi‑task extension of ProtBERT designed for classifying intrinsically disordered proteins (IDPs) in low‑data settings. It combines Dynamic Window Masking, a Multi‑Scale 1D Convolutional classifier, and auxiliary biochemistry‑informed objectives to jointly optimize masked language modeling and domain‑specific tasks. In experiments on phosphorylation site prediction and protein compaction prediction, MT‑ProtBERT outperforms the RNN‑based IDP model PARROT across all limited‑data tasks.
By Jian Sun, Kingshuk Ghosh, Lilianna Houston, Mohammad H. Mahoor
StabilityArc is a method that decodes protein sequence embeddings into generalizable stability landscapes. It uses a shared RoPE transformer to map frozen ESMC-600M residue representations into an Lx20 matrix of substitution effects, with a symmetric, contact-aware residual to predict epistasis. In extensive leave-one-protein-out tests on 134,794 ProteinGym variants, StabilityArc achieves a Spearman correlation of 0.7134, surpassing the best zero‑shot baseline, and further improves Kermut’s performance when used as a prior.
By Aaron L. Feller, Andrew D. Ellington, Claus O. Wilke
FLaG (Frequency‑Domain Latent‑attention Gated Pooling) is a plug‑in token‑aggregation module that transforms encoder outputs into the Fourier domain, summarizes spectral tokens with learnable latent queries, applies a sample‑conditioned channel gate, and reconstructs modulated token representations for downstream pooling. The method is evaluated on antimicrobial peptide activity prediction, CIFAR‑10/100 image classification, and several RoBERTa language tasks, achieving state‑of‑the‑art performance on most metrics. Analyses show that FLaG emphasizes low‑frequency components while selectively amplifying high‑frequency signals in later layers, providing a transferable frequency‑domain bias across protein, visual, and textual representations.
By Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Ruochi Zhang, Fengfeng Zhou
arXiv:2608. 01924v2 Announce Type: replace-cross Abstract: Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe class imbalance trait and the complex of immunogenicity features.
By Zhiyin An, Yuenan Hou, Shumeng Duan, Yiming Zhou, Yuanting Zheng, Leming Shi
The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.
By Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary, Etienne Th\'evenot, Florence d'Alch\'e-Buc
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE...
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.
By Raunak Kumar, Anuj Pal, Dhruvi Solanki, Parikshit Pareek, Juhi Singh, Jitin Singla
arXiv:2608.21367v1 Announce Type: cross
Abstract: Protein-peptide interactions are central to cellular regulation and peptide-based drug discovery, yet existing computational methods mainly focus on...
By Hao Qian, Shikui Tu, Lei Xu