FLaG (Frequency‑Domain Latent‑attention Gated Pooling) is a plug‑in token‑aggregation module that transforms encoder outputs into the Fourier domain, summarizes spectral tokens with learnable latent queries, applies a sample‑conditioned channel gate, and reconstructs modulated token representations for downstream pooling. The method is evaluated on antimicrobial peptide activity prediction, CIFAR‑10/100 image classification, and several RoBERTa language tasks, achieving state‑of‑the‑art performance on most metrics. Analyses show that FLaG emphasizes low‑frequency components while selectively amplifying high‑frequency signals in later layers, providing a transferable frequency‑domain bias across protein, visual, and textual representations.
By Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Ruochi Zhang, Fengfeng Zhou
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE...
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.
By Raunak Kumar, Anuj Pal, Dhruvi Solanki, Parikshit Pareek, Juhi Singh, Jitin Singla
Freeze, Diffuse, Decode (FDD) is a diffusion-based framework that adapts pre‑trained transformer embeddings to downstream tasks while preserving their geometric structure. By propagating supervised signals along the intrinsic manifold of frozen embeddings, FDD produces low‑dimensional, predictive, and interpretable representations. In antimicrobial peptide design, these representations support property prediction, retrieval, and latent‑space interpolation.
By Pankhil Gawade, Adam Izdebski, Myriam Lizotte, Kevin R. Moon, Jake S. Rhodes, Guy Wolf, Ewa Szczurek
arXiv:2511.23120v2 Announce Type: replace
Abstract: Pretrained transformers provide rich, general-purpose embeddings, which are transferred to downstream tasks. However, current transfer strategies:...
By Pankhil Gawade, Adam Izdebski, Myriam Lizotte, Kevin R. Moon, Jake S. Rhodes, Guy Wolf, Ewa Szczurek
The paper introduces DPTM‑DT, a dual‑pretrained Transformer framework that integrates GROVER molecular graph embeddings, ESM protein language‑model embeddings, and CTD physicochemical descriptors for drug‑target prediction. It employs bidirectional cross‑modal attention to share drug‑target information and uses a single pair representation for continuous affinity regression, high‑affinity binary classification, and six‑level affinity classification. Experiments on Davis and KIBA datasets show that DPTM‑DT outperforms existing methods across regression, binary, and multiclass tasks, with ablation studies confirming the contributions of dual target representation, gated fusion, and cross‑modal attention.
By Ge Kong
arXiv:2607. 14070v1 Announce Type: cross Abstract: Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored.
By Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jan\v{c}ovi\v{c}ov\'a, Kri\v{z}an Jurinovi\'c, Vanessa Smilansky
arXiv:2602. 22822v3 Announce Type: replace Abstract: Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice.
By Yunhua Zhong, Yixuan Tang, Yifan Li, Pan Liu, Zhiwen Yang, Jie Yang, Jun Xia
arXiv:2504. 17247v3 Announce Type: replace Abstract: Deep learning-based antimicrobial peptide (AMP) discovery faces critical challenges such as limited controllability, lack of representations that efficiently model antimicrobial properties, and low experimental hit rates.
By Diogo Soares, Leon Hetzel, Paulina Szymczak, Marcelo Der Torossian Torres, Johanna Sommer, Cesar de la Fuente-Nunez, Fabian Theis, Stephan G\"unnemann, Ewa Szczurek
The paper presents a deep learning approach that converts nanopore resistive pulse signals into scaleograms using continuous wavelet transforms, enabling the classification of peptides as an image‑classification problem. On a dataset of 42 peptides, the method achieves an 82% macro‑averaged accuracy, outperforming previous descriptor‑based techniques by 8.6 percentage points. The models also remain accurate after significant weight pruning and 8‑bit quantization, making them suitable for deployment on embedded sensing hardware.
By Julian Ho{\ss}bach, Samuel Tovey, Sandro Kuppel, Tobias Ensslen, Jan C. Behrends, Christian Holm
arXiv:2608.30175v1 Announce Type: new
Abstract: Peptide-protein affinity models are often evaluated with a single data split, obscuring whether they interpolate among measurements for observed target...
By Jiaxin Tian, Darren An, Jun Li
GyroNovo is a new framework for de novo peptide sequencing that improves fragment imputation by guiding the process with decoder errors observed during training. It introduces mass-aware attention using rotary embeddings to encode pairwise mass differences between spectral peaks, and creates easy and hard augmented views of spectra to train the decoder under varying corruption levels. Experiments on NovoBench demonstrate significant gains, with about 9 percentage points higher peptide-level precision and 7 percentage points higher amino-acid-level precision compared to the state-of-the-art baseline.
By Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed