arXiv Machine Learning By Hunter Heidenreich

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES

Read the original on arXiv Machine Learning →

arXiv:2607. 05691v1 Announce Type: cross Abstract: Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

The paper audits 22 frontier language models on 12 molecular property regression benchmarks to assess verbatim retrieval of published values. It finds widespread but benchmark‑specific retrieval, with over 50% of models retrieving exact values on five datasets and isolated occurrences on others. Experiments at different reasoning levels show that higher reasoning increases retrieval flags, and attempts to interrupt retrieval reveal that top models can still recognize transformed SMILES and original labels. Suppressing retrieval reduces prediction error variance, indicating that predictive performance is not solely due to memorized values.

By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler
arXiv AI
Sep 18

Chunk Twice, Embed Once: A Systematic Study of Segmentation and Representation Trade-offs in Chemistry-Aware Retrieval-Augmented Generation

The study investigates how document segmentation and chunk representation affect retrieval-augmented generation (RAG) for chemistry texts. Using the ChemQuests corpus, the authors benchmark 41 embedding models and evaluate them across five chunking strategies, seven chunk sizes, and various overlap settings. They find that embedding choice has the largest impact, with models like E5, BGE, and Nomic performing best, and recommend medium-to-large chunks with fixed-token, recursive-token, or hierarchical-section chunking and low overlap for effective chemistry-aware RAG.

By Mahmoud Amiri, Thomas Bocklitz