arXiv Machine Learning By Tirtharaj Dash, Gunja Sachdeva

$p$-adic Bi-Filtrations for Topological Machine Learning on Genomic Sequences

Read the original on arXiv Machine Learning →

arXiv:2606. 06117v1 Announce Type: cross Abstract: We introduce pVR, a topological machine learning framework for alignment-free genomic sequence classification that combines $p$-adic numbers with topological data analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jun 4

$p$-adic Bi-Filtrations for Topological Machine Learning on Genomic Sequences

We introduce pVR, a topological machine learning framework for alignment-free genomic sequence classification that combines $p$-adic numbers with topological data analysis. Each DNA sequence is encoded along two complementary axes: a $p$-adic distance on $k$-mer prefixes, which captures hierarchical positional structure, and a compositional $L_1$ distance on $k$-mer frequencies, which captures local sequence content.

arXiv Machine Learning
Sep 11

When do cheap embeddings beat protein language models? A theoretically-grounded hashing sketch for biological sequence classification

The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.

By Sarwan Ali, Taslim Murad, Imdadullah Khan, Safi Faizullah
arXiv Machine Learning
Jun 16

Learning Topological Representations for Molecular Dynamics

arXiv:2606. 14737v1 Announce Type: cross Abstract: Molecular dynamics (MD) simulations generate trajectories in a high-dimensional configuration space whose analysis critically depends on molecular descriptors, typically handcrafted observables or learned kinetic embeddings.

By Dominik Geng, Florian Graf, Martin Uray, Roland Kwitt
arXiv Machine Learning
Jun 2

Prototype Selection Using Topological Data Analysis

arXiv:2511. 04873v2 Announce Type: replace-cross Abstract: Prototype selection methods compress a training set, but the existing taxonomy of condensation, edition, hybrid, competence-based, optimization-based, and clustering-based families does not include methods that operate on the multi-scale topological structure of the data.

By Jordan Eckert, Elvan Ceyhan, Henry Schenck
arXiv Machine Learning
Sep 21

COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

COMPLEX is a closed‑form, training‑free embedding for multiparameter persistence modules that provides both an upper and a lower Lipschitz bound, enabling faithful feature representations. By slicing modules along a near‑diagonal net and embedding each slice with the certified PLACE/PALACE landmark map, the method guarantees that separated modules remain separated in the embedding. On Orbit benchmarks and molecular graph tasks, COMPLEX achieves state‑of‑the‑art accuracy, outperforming existing landmark, transformer, and graph‑based approaches.

By Sushovan Majhi, Atish Mitra, \v{Z}iga Virk, Pramita Bagchi