arXiv Machine Learning

Structured Gaussian Processes for Uncertainty-Aware Classification of High-Dimensional, Small-Sampled Omics Data

arXiv:2607. 02103v1 Announce Type: cross Abstract: Classifying heterogeneous omics data remains a fundamental challenge in computational biology, particularly in high-dimensional, small-sample settings where nonlinear interactions dominate and class imbalance further complicates reliable prediction of minority phenotypes.

arXiv Machine Learning
Jun 30

Friend or Foe

arXiv:2509. 00123v2 Announce Type: replace-cross Abstract: A fundamental challenge in microbial ecology is determining whether bacteria compete or cooperate in different environmental conditions.

By Oleksandr Cherednichenko, Josephine Solowiej-Wedderburn, Laura M. Carroll, Eric Libby
arXiv Machine Learning
Sep 23

Hierarchical Sparse Bayesian Multitask Learning for Disease Prediction in Pooled Microbiome Studies

This paper introduces a hierarchical Bayesian multitask learning model that assumes a shared sparsity structure across different binary classification tasks. The authors develop a variational inference algorithm for efficient posterior approximation and evaluate the method on synthetic data and pooled microbiome studies. Results show superior support recovery in synthetic experiments and robust, well‑calibrated predictions with informative taxa selection in microbiome classification.

By Haonan Zhu, Andre R. Goncalves, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Mart\'i, Car Reen Kok, Monica K. Borucki, Nisha J. Mulakken, James B. Thissen, Crystal Jaing, Alfred Hero, Nicholas A. Be
arXiv AI
Aug 11

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

arXiv:2608. 08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identification and antimicrobial resistance prediction.

By Alejandro L. Garc\'ia-Navarro, Carlos Sevilla-Salcedo, Bel\'en Rodr\'iguez-S\'anchez, Vanessa G\'omez-Verdejo
arXiv Machine Learning
Aug 28

MODIS: Multi-Omics Data Integration for Small and unpaired datasets

MODIS is a semi‑supervised framework for integrating multi‑omics data that are often unpaired, partially labeled, and scarce, such as in rare disease studies. It trains on a large reference database and a small target dataset simultaneously, using diagonal integration and class‑label alignment to handle class imbalance. The architecture combines variational auto‑encoders, a class classifier, and an adversarially trained modality classifier, with a regularized relativistic GAN loss for stable training, and demonstrates high accuracy on synthetic data and the TCGA cancer dataset.

By Daniel Lepe-Soltero, Thierry Arti\`eres, Ana\"is Baudot, Paul Villoutreix
arXiv AI
Sep 10

The Accuracy Paradox: Empirical Diagnostic of Default Decision Thresholds in Multi-Label Enzyme Commission Prediction [With Code]

The study evaluates the use of default decision thresholds (t=0.50) in multi‑label enzyme commission (EC) number prediction across 14,096 compounds and six EC classes. It finds a high mean accuracy of 77.16% but low macro F1 (0.3976) and macro recall (0.3872), indicating severe class‑imbalance issues: majority classes are over‑predicted while minority classes, especially EC6, have zero recall despite reasonable ROC‑AUC. The authors recommend target‑specific threshold tuning and conformal calibration as post‑processing safeguards to expose and correct these hidden errors.

By Bilal Ahmad, Rajed Mehmood