Beyond Attention: Signed Integrated Gradients Attribution in a BiomeGPT-Style Microbiome Transformer
arXiv:2608. 06486v1 Announce Type: new Abstract: In a feature-tokenized transformer (arXiv:2106.
arXiv:2606. 24995v1 Announce Type: new Abstract: Tabular foundation models (TFMs) achieve strong performance on microbiome abundance data, yet their robustness under realistic distribution shift remains poorly characterized.
arXiv:2608. 06486v1 Announce Type: new Abstract: In a feature-tokenized transformer (arXiv:2106.
arXiv:2605.28868v2 Announce Type: replace-cross Abstract: Metagenomic taxonomic annotation is essential for interpreting complex microbial communities, yet reliable annotation remains challenging und...
The study demonstrates that a simple, sequence-only approach using 330 interpretable descriptors and the TabPFN tabular foundation model can outperform complex multimodal deep learning methods for multi-label antimicrobial peptide activity prediction. On the ESCAPE benchmark (82,359 peptides, five labels), a label‑powerset TabPFN model achieved a mean average precision of 77.8%, surpassing the previous best of 72.1%. The approach also shows that predicted structure is unnecessary, that a small set of global physicochemical scalars can recover most performance, and that modeling label dependence benefits rare activities and informs assay prioritization.
arXiv:2607. 02103v1 Announce Type: cross Abstract: Classifying heterogeneous omics data remains a fundamental challenge in computational biology, particularly in high-dimensional, small-sample settings where nonlinear interactions dominate and class imbalance further complicates reliable prediction of minority phenotypes.
arXiv:2607. 25497v1 Announce Type: cross Abstract: Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centres.
Antimicrobial peptides (AMPs) often act against multiple pathogen classes, making multi-label activity prediction a more realistic screening target than binary antimicrobial classification. The ESCAPE...
arXiv:2607. 25497v2 Announce Type: replace-cross Abstract: Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions.
arXiv:2610.01143v1 Announce Type: cross Abstract: Despite strong mean accuracy, tabular foundation models (TFMs) can perform poorly on underrepresented groups under subpopulation shift, where group p...
The paper introduces SAGE, a Sampling‑Aware Global Evaluation benchmark for species distribution modeling that uses GBIF records for training and sPlotOpen vegetation plots for presence‑absence evaluation across 5,771 plant species. It groups species by sampling effort and relative prevalence to assess how well single‑species and multi‑species deep‑learning SDMs perform under different data conditions. The study finds that Random Forests and DeepSDMs perform best overall, with DeepSDMs excelling for infrequently recorded species only when bias‑correction techniques are applied.
arXiv:2607. 14070v1 Announce Type: cross Abstract: Genomic foundation models such as Evo 2 learn rich sequence representations, but their value for biosecurity screening is largely unexplored.
This paper introduces a hierarchical Bayesian multitask learning model that assumes a shared sparsity structure across different binary classification tasks. The authors develop a variational inference algorithm for efficient posterior approximation and evaluate the method on synthetic data and pooled microbiome studies. Results show superior support recovery in synthetic experiments and robust, well‑calibrated predictions with informative taxa selection in microbiome classification.
arXiv:2605. 28418v3 Announce Type: replace Abstract: With the rise of tabular foundation models alongside traditional models still performing well on many tasks, choosing the right model for a tabular dataset remains difficult.