Sharpness-Aware Minimization (SAM) is applied to improve the generalization of machine learning models for classifying bacterial Raman spectra, a technique that could enable rapid, portable diagnostics for antimicrobial resistance. The study shows that SAM can increase classification accuracy by up to 10.5% on a single data split and by an average of 2.7% across multiple splits compared to the traditional Adam optimizer. These gains demonstrate SAM’s potential to enhance the clinical utility of AI-powered Raman spectroscopy tools.
By Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari, Hossein Mobahi, Loza F. Tadesse
Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring. Predictive Raman datasets often contain few labelled spectra and thousands of ordered wavenumbers, with informative variation within bands and across distant spectral regions.
arXiv:2608. 02157v1 Announce Type: new Abstract: Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring.
By Xingyu Pan, Huan Wang, Jinjia Guo, Zhenlin Zhao, Siming Dong, Jixi Lu
arXiv:2606. 27096v1 Announce Type: new Abstract: Transformer-based models have recently attracted increasing attention for Raman spectral classification.
By Jamile Mohammad Jafari, Thomas Bocklitz
arXiv:2607. 10196v1 Announce Type: new Abstract: Access to sufficiently large biomedical datasets remains a major obstacle for machine learning in Raman spectroscopy-based diagnostics.
By Andrei Iu\c{s}an, Iulian Vasile, Daria Voiculescu, Ion Petre, Andrei P\u{a}un, Bogdan Oancea, Mihaela P\u{a}un
arXiv:2608. 13341v1 Announce Type: cross Abstract: Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging.
By Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled training sets, and transfer poorly across analytical objectives and experimental datasets.
The paper presents a Raman spectroscopy and machine‑learning framework that uses t‑SNE, K‑means, Decision Trees, and NNLS‑based spectral decomposition to authenticate edible oils. In pure oils, Decision Trees achieved 100% accuracy using only four Raman variables out of 1866 features, while in matrix‑containing samples, NNLS‑based PI‑AI improved classification to about 85–86% with only four to five key variables. The approach yields highly compact, interpretable spectral representations that enable accurate oil identification with minimal data footprint.
By Amrita Shaw, Chandrasekar S. N., Sai Muthukumar V., Jhinuk Gupta, Deepak L. N. Kallepalli
arXiv:2606. 26520v1 Announce Type: cross Abstract: Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene.
By Johnny Peng, Thanh Tung Khuat, Ellen Otte, Katarzyna Musial, Bogdan Gabrys
arXiv:2606. 30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets.
By Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart, Fabian Pauly
The paper introduces MSAlign, a lightweight model that aligns frozen foundation models for mass spectra (DreaMS) and molecules (MolDeBERTa) to improve metabolite identification from MS/MS spectra. It presents a unified framework for representation alignment and contrastive learning, demonstrates that a score fusion strategy further boosts performance at minimal cost, and addresses evaluation challenges by quantifying distribution shift in data splitting strategies. All resources, including datasets, splits, and code, are publicly released to promote reproducible research.
By Paul Krzakala, Gabriel Melo, Camille Lan\c{c}on, Charlotte Laclau, R\'emi Flamary, Etienne Th\'evenot, Florence d'Alch\'e-Buc
arXiv:2602. 22822v3 Announce Type: replace Abstract: Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice.
By Yunhua Zhong, Yixuan Tang, Yifan Li, Pan Liu, Zhiwen Yang, Jie Yang, Jun Xia