arXiv Machine Learning

Machine Learning for Biomedical Raman Spectroscopy: From Spectral Acquisition to Clinical Translation

arXiv:2606. 14169v1 Announce Type: new Abstract: Raman spectroscopy provides label-free, chemically specific characterization of biological systems and has become an important tool for cancer diagnosis, molecular subtyping, microbiological identification, and intraoperative decision support.

arXiv Machine Learning
Sep 18

Sharpness-Aware Minimization (SAM) Improves Classification Accuracy of Bacterial Raman Spectral Data Enabling Portable Diagnostics

Sharpness-Aware Minimization (SAM) is applied to improve the generalization of machine learning models for classifying bacterial Raman spectra, a technique that could enable rapid, portable diagnostics for antimicrobial resistance. The study shows that SAM can increase classification accuracy by up to 10.5% on a single data split and by an average of 2.7% across multiple splits compared to the traditional Adam optimizer. These gains demonstrate SAM’s potential to enhance the clinical utility of AI-powered Raman spectroscopy tools.

By Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari, Hossein Mobahi, Loza F. Tadesse
arXiv AI
Aug 14

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

arXiv:2608. 13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging.

By Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Hugging Face Trending Papers
Aug 13

Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples

Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled training sets, and transfer poorly across analytical objectives and experimental datasets.

arXiv AI
Aug 24

Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

The paper presents a Raman spectroscopy and machine‑learning framework that uses t‑SNE, K‑means, Decision Trees, and NNLS‑based spectral decomposition to authenticate edible oils. In pure oils, Decision Trees achieved 100% accuracy using only four Raman variables out of 1866 features, while in matrix‑containing samples, NNLS‑based PI‑AI improved classification to about 85–86% with only four to five key variables. The approach yields highly compact, interpretable spectral representations that enable accurate oil identification with minimal data footprint.

By Amrita Shaw, Chandrasekar S. N., Sai Muthukumar V., Jhinuk Gupta, Deepak L. N. Kallepalli
arXiv AI
Jun 26

Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting

arXiv:2606. 26520v1 Announce Type: cross Abstract: Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene.

By Johnny Peng, Thanh Tung Khuat, Ellen Otte, Katarzyna Musial, Bogdan Gabrys
arXiv Machine Learning
5d ago

MSAlign: Aligning Molecule and Mass Spectra representations for Metabolite Identification

The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.

By Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary, Etienne Th\'evenot, Florence d'Alch\'e-Buc