arXiv Machine Learning By Stylianos Kampakis, Fabio Rovai

What an odour descriptor corpus can and cannot measure: valence, attenuation, and the ceiling of the public record

Read the original on arXiv Machine Learning →

The study evaluates the reliability of shared odor descriptor words across four public corpora from Pyrfume, finding substantial disagreement (I² = 80 %) and limited agreement on descriptor application (median tetrachoric = 0.795, κ = 0.413). Only a fraction of the achievable variance in odor perception is captured by current models and descriptor sets, with valence emerging as the primary missing component. Even with extensive model capacity and merged corpora, the gap remains, indicating that valence must be measured directly to improve machine olfaction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 17

English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck

The paper reports that in English all‑words word sense disambiguation (WSD), the scarcity of high‑quality labels—not the models—has become the limiting factor. The authors introduce lexEN, a human‑adjudicated correction layer over the Maru2022 ALL_NEW benchmark, and SenseBench, a living leaderboard for LLM WSD evaluation. They show that frontier large language models reach about 95 % accuracy on lexEN‑v1, that relabeling corpora with these models improves downstream systems, and that fine‑grained WordNet senses are often ill‑posed, with coarsening improving both annotator agreement and model performance. "whyItMatters":"The study highlights that improving label quality and managing annotation costs are now the critical challenges for advancing WSD performance, as model accuracy is already near its theoretical ceiling."

By Vassili Philippov, Amro Salman, Dmitrii Andreev, Penny Hands, Emil Kaiumov, Pavel Katunin, Anton Nikolaev
arXiv Computer Vision
Aug 25

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

The paper introduces VERDICT, a method for validating optical chemical structure recognition (OCSR) outputs by leveraging agreement among multiple recognizers rather than pixel‑space re‑rendering. On 263 ACS journal images, agreement achieved an AUROC of 0.916, far surpassing the 0.547 AUROC of re‑rendering similarity. VERDICT was applied to PMC Open Access, yielding over 6,000 high‑precision structure labels, and is integrated into SES AI’s Molecular Universe platform for image‑based molecular search.

By Yani Guan, Dengpan Dong, Shuang Luo, Zi Wei, Joah Han, Dan Hannah, Yumin Zhang, Qichao Hu, Kang Xu
arXiv Machine Learning
Sep 11

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

The paper investigates whether language models can identify sentences from their training data by using exact duplication counts from publicly released corpora for two model families, OLMo‑2 and Pythia. It finds that for typical duplication levels, models show only a weak trace of exposure, with a rank correlation near –0.08, and that strong signals only appear when a sentence appears roughly a thousand times, at which point fame rather than memory dominates. The study also demonstrates that common membership tests can be misleading, as changing a single word does not alter the model’s preference, and that controlling for register can significantly improve detector performance.

By Arman Nik Khah
arXiv AI
Sep 7

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

The paper audits 22 frontier language models on 12 molecular property regression benchmarks to assess verbatim retrieval of published values. It finds widespread but benchmark‑specific retrieval, with over 50% of models retrieving exact values on five datasets and isolated occurrences on others. Experiments at different reasoning levels show that higher reasoning increases retrieval flags, and attempts to interrupt retrieval reveal that top models can still recognize transformed SMILES and original labels. Suppressing retrieval reduces prediction error variance, indicating that predictive performance is not solely due to memorized values.

By Matthias Busch, Marius Tacke, Sviatlana V. Lamaka, Mikhail L. Zheludkevich, Christian J. Cyron, Roland C. Aydin, Christian Feiler