arXiv AI

Fusing Sequence Motifs and Pan-Genomic Features: Antimicrobial Resistance Prediction using an Explainable Lightweight 1D CNN-XGBoost Ensemble

arXiv AI
Aug 11

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

arXiv:2608. 08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction.

By Alejandro L. Garc\'ia-Navarro, Carlos Sevilla-Salcedo, Bel\'en Rodr\'iguez-S\'anchez, Vanessa G\'omez-Verdejo
arXiv Machine Learning
Jun 30

Friend or Foe

arXiv:2509. 00123v2 Announce Type: replace-cross Abstract: A fundamental challenge in microbial ecology is determining whether bacteria compete or cooperate in different environmental conditions.

By Oleksandr Cherednichenko, Josephine Solowiej-Wedderburn, Laura M. Carroll, Eric Libby
arXiv Machine Learning
Sep 1

Coarse composition suffices: tabular in-context learning for multi-activity antimicrobial peptide profiling

The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.

By Raunak Kumar, Anuj Pal, Dhruvi Solanki, Parikshit Pareek, Juhi Singh, Jitin Singla
arXiv Machine Learning
Aug 28

MODIS: Multi-Omics Data Integration for Small and unpaired datasets

MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.

By Daniel Lepe-Soltero, Thierry Arti\`eres, Ana\"is Baudot, Paul Villoutreix