arXiv:2606. 24995v1 Announce Type: new Abstract: Tabular foundation models (TFMs) achieve strong performance on microbiome abundance data, yet their robustness under realistic distribution shift remains poorly characterized.
By Giulia Perciballi, Ahmad Fall, Federica Granese, Edi Prifti, Jean-Daniel Zucker
arXiv:2607. 14070v1 Announce Type: cross Abstract: Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored.
By Jeremy Guntoro, Alexander Dack, Dylan Danno, Michaela Jan\v{c}ovi\v{c}ov\'a, Kri\v{z}an Jurinovi\'c, Vanessa Smilansky
arXiv:2509.23552v2 Announce Type: replace-cross
Abstract: Antimicrobial Resistance (AMR) is a rapidly escalating global health crisis. While genomic sequencing enables rapid prediction of resistance...
By Md. Saiful Bari Siddiqui, Nowshin Tarannum
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.
By Raunak Kumar, Anuj Pal, Dhruvi Solanki, Parikshit Pareek, Juhi Singh, Jitin Singla
arXiv:2608. 02684v1 Announce Type: cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse.
By Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji
arXiv:2607. 27258v1 Announce Type: cross Abstract: Plant biosynthetic gene clusters (BGCs) encode specialized-metabolite pathways, yet curated plant BGC labels remain scarce, hindering supervised discovery at genome scale.
By Yuhan Zhao, Nidhi Grover, Zhishan Guo, Ning Sui
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE...
arXiv:2605.28868v2 Announce Type: replace-cross
Abstract: Metagenomic taxonomic annotation is essential for interpreting complex microbial communities, yet reliable annotation remains challenging und...
By Rongye Ye, Lun Li, Zheng Luo, Yiran Zhan, Zhang Zhang, Shuhui Song
FLaG (Frequency‑Domain Latent‑attention Gated Pooling) is a plug‑in token‑aggregation module that transforms encoder outputs into the Fourier domain, summarizes spectral tokens with learnable latent queries, applies a sample‑conditioned channel gate, and reconstructs modulated token representations for downstream pooling. The method is evaluated on antimicrobial peptide activity prediction, CIFAR‑10/100 image classification, and several RoBERTa language tasks, achieving state‑of‑the‑art performance on most metrics. Analyses show that FLaG emphasizes low‑frequency components while selectively amplifying high‑frequency signals in later layers, providing a transferable frequency‑domain bias across protein, visual, and textual representations.
By Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Qiuchen Zhao, Lan Huang, Ruochi Zhang, Fengfeng Zhou
arXiv:2606. 07686v1 Announce Type: cross Abstract: Physics-Informed Neural Network (PINN) is a way of including knowledge in the form of equations in Machine Learning methods.
By Ravisha Rupasinghe, Rajith Vidanaarachchi, Asela Hevapathige, Sachith Seneviratne, Sen-Lin Tang, Saman Halgamuge
arXiv:2606. 01042v1 Announce Type: cross Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions.
By Xinyu Yuan, Xixian Liu, Jianan Zhao, Yashi Zhang, Hongyu Guo, Jian Tang
arXiv:2606. 08191v1 Announce Type: cross Abstract: Token aggregation is a common bottleneck in models that map token representations to sample-level predictions, yet most pooling methods operate only in the original token domain.
By Kewei Li, Rongying Zhang, Xueli Wang, Xiwen Gong, Zhongjian Wang, Lan Huang, Ruochi Zhang, Fengfeng Zhou