arXiv:2606. 14169v1 Announce Type: new Abstract: Raman spectroscopy provides label-free, chemically specific characterization of biological systems and has become an important tool for cancer diagnosis, molecular subtyping, microbiological identification, and intraoperative decision support.
By Bogdan Oancea, Ana Maria Seciu-Grama, Nicoleta Siminea, Laura Mihaela Stefan, Alice Stoica, Joel Sjoberg, Marian Necula, Ana-Maria Prelipcean, Corneliu Ovidiu Vrancianu, Eduard Milea, Andrei P\u{a}un, Ion Petre, Mihaela P\u{a}un
arXiv:2608. 02157v1 Announce Type: new Abstract: Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring.
By Xingyu Pan, Huan Wang, Jinjia Guo, Zhenlin Zhao, Siming Dong, Jixi Lu
Raman spectroscopy enables non-destructive, label-free molecular characterization across materials science, biomedicine and process monitoring. Predictive Raman datasets often contain few labelled spectra and thousands of ordered wavenumbers, with informative variation within bands and across distant spectral regions.
arXiv:2610.02224v1 Announce Type: new
Abstract: Rapid and reliable identification of pharmaceutical residues is important for safeguarding public health, ensuring food safety, and enabling practical...
By Quach Thi Thai Binh, Ton Nu Quynh Trang, Thang B. Phan, Vu Thi Hanh Thu, Nguyen Tuan Hung
Sharpness-Aware Minimization (SAM) is applied to improve the generalization of machine learning models for classifying bacterial Raman spectra, a technique that could enable rapid, portable diagnostics for antimicrobial resistance. The study shows that SAM can increase classification accuracy by up to 10.5% on a single data split and by an average of 2.7% across multiple splits compared to the traditional Adam optimizer. These gains demonstrate SAM’s potential to enhance the clinical utility of AI-powered Raman spectroscopy tools.
By Kaitlin Zareno, Jarett Dewbury, Siamak K. Sorooshyari, Hossein Mobahi, Loza F. Tadesse
BOOM is a new benchmark for evaluating out‑of‑distribution (OOD) molecular property predictions in machine learning. It provides chemically‑informed tests across common property prediction tasks and assesses over 150 model‑task combinations. The study shows that current models, including chemical foundation models, struggle to generalize OOD, with the best model still exhibiting three times higher error than in‑distribution predictions.
By Evan R. Antoniuk, Shehtab Zaman, Tal Ben-Nun, Peggy Li, James Diffenderfer, Busra Sahin, Obadiah Smolenski, Everett Grethel, Tim Hsu, Anna M. Hiszpanski, Kenneth Chiu, Bhavya Kailkhura, Brian Van Essen
arXiv:2602. 22822v3 Announce Type: replace Abstract: Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction still remain difficult to evaluate and deploy in practice.
By Yunhua Zhong, Yixuan Tang, Yifan Li, Pan Liu, Zhiwen Yang, Jie Yang, Jun Xia
arXiv:2607. 10196v1 Announce Type: new Abstract: Access to sufficiently large biomedical datasets remains a major obstacle for machine learning in Raman spectroscopy-based diagnostics.
By Andrei Iu\c{s}an, Iulian Vasile, Daria Voiculescu, Ion Petre, Andrei P\u{a}un, Bogdan Oancea, Mihaela P\u{a}un
arXiv:2606. 30961v1 Announce Type: cross Abstract: Advances in deep learning architectures and representations have enabled ML-driven chemical property prediction, but state-of-the-art (SOTA) models have remained largely confined to independent codebases and lack support for diverse chemical species.
By Jacob W. Toney, Samir Darouich, Yiran Wang, Aaron G. Garrison, Johannes K\"astner, Heather J. Kulik
TabBench-Bio is a living, interactive benchmark that evaluates machine learning models on 43 high‑dimensional biomedical tables, covering multiple domains. Using a shared cross‑validation protocol, the benchmark compares classical estimators, neural networks, and tabular foundation models across 28 feature‑by‑sample operating points, with RealTabPFN v2.5 achieving the highest performance at the reference cell of 10,000 features and 100 training samples. The benchmark provides reproducible results, fold‑level predictions, and invites community contributions to expand its dataset collection.
By Jules Kreuer, Sofiane Ouaari, Julia Hellmig, Julius Braitinger, Nico Pfeifer
arXiv:2606. 30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets.
By Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart, Fabian Pauly
arXiv:2606. 26520v1 Announce Type: cross Abstract: Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process parameters drift over days, and an off-specification trend is often confirmed too late to intervene.
By Johnny Peng, Thanh Tung Khuat, Ellen Otte, Katarzyna Musial, Bogdan Gabrys