arXiv Machine Learning

Knowledge-guided Pattern Discovery via Coupled Tensor Factorizations

arXiv:2608. 13234v1 Announce Type: new Abstract: In order to understand complex systems such as the human metabolome or human brain, different sensing technologies are used, generating complex data.

arXiv Machine Learning
Sep 4

Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing

The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both runtime and root‑sum‑of‑squares error for drug effect prediction.

By Maryam Bagherian, Albert Hung, Ivo Dinov, Joshua Welch
Hugging Face Trending Papers
Sep 2

Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing

The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both run‑time and root‑sum‑of‑errors accuracy for predicting drug effects.

arXiv Machine Learning
5d ago

MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification

The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.

By Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary, Etienne Th\'evenot, Florence d'Alch\'e-Buc
arXiv Statistics ML
Aug 25

Neuro-Causal Factor Analysis

Neuro-Causal Factor Analysis (NCFA) reimagines traditional factor analysis by integrating causal structure learning and deep generative modeling. The method learns a directed graph linking latent and observed variables, then trains a deep generative model that respects the graph’s Markov factorization. Experiments on synthetic and real datasets show NCFA achieves lower reconstruction error than standard FA and better latent distribution recovery than a variational autoencoder, while offering a sparser architecture, reduced complexity, and causal interpretability.

By Alex Markham, Mingyu Liu, Bryon Aragam, Liam Solus
arXiv Machine Learning
Aug 20

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Monroe is a new molecular foundation model that improves upon existing models by pre‑training on over 81 million molecules from the PM6 quantum chemistry dataset, enhancing stereochemistry representation, and introducing novel training losses such as conformer denoising and embedding decorrelation. It also incorporates a prior‑data‑fitted model (TabPFN) for downstream in‑context prediction and demonstrates superior performance on Polaris benchmarks and activity cliff tests. Ablation studies show that the PFN‑based downstream approach can upgrade other models, producing state‑of‑the‑art variants MiniMol_PFN and CheMeleon_PFN.

By Blazej Banaszewski, Andrew W. Fitzgibbon
Hugging Face Trending Papers
Aug 20

Orthogonal JEPA: Factorized Predictive States for Latent World Models

Orthogonal JEPA introduces a latent world‑modeling framework that factorizes predictive states into orthogonal components. By learning basis matrices and dedicated prediction branches, the method reduces redundancy and improves gradient signals for less dominant predictive structures. The factorized states can be synthesized into complete latent representations for downstream tasks such as decoding, planning, or autoregressive rollout, and are evaluated across vision, biology, health, control, and molecular dynamics domains.