arXiv AI By Anna T. Thomas, Sohum Patnaik, Caroline Cotto, Benjamin Sanchez-Lengeling

TasteBench: Multimodal Benchmark for Sensory Prediction, from Molecules to Sustainable Foods

Read the original on arXiv AI →

TasteBench is a multimodal benchmark designed to accelerate sustainable protein discovery by providing computational proxies for sensory prediction. It includes a food-level ranking task based on over 21,000 human evaluations of 215 plant-based foods across 24 categories, and a molecular-level taste classification task covering 15,000 flavor molecules. The benchmark offers baseline models, characterizes inter-rater agreement and reliability limits, and demonstrates that the best model achieves pairwise accuracy comparable to individual human panelists.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Scientific Discovery under Validation Congestion via Multi-Fidelity Pairwise Rankings

The paper introduces PRISMS, a framework that uses expert pairwise rankings of varying fidelity to curate scientific designs without relying on data-intensive regression models. By escalating queries from lower- to higher-fidelity rankers based on Fisher-information, PRISMS improves discovery recall and reduces the number of screening rounds compared to regression-only and non‑escalated ranking methods. In optimization tasks, PRISMS outperforms Bayesian optimization by achieving higher hypervolume.

By Kevin Tirta Wijaya, Alston Lo, Michael Sun, Wojciech Matusik, Vahid Babaei
arXiv Machine Learning
Sep 29

Interpretable-by-Design Descriptor Portfolios Match a 2048-Dimensional Foundation Embedding on Low-Data Molecular Assays

The study evaluates whether a portfolio of compact, semantically named descriptor blocks can match the performance of a 2048‑dimensional CheMeleon embedding in low‑data molecular assays. Using a fixed 11‑dimensional physicochemical base and greedily adding provenance‑screened blocks, the portfolio achieves a mean test AUC of 0.762 across nine ADME/Tox assays, comparable to CheMeleon’s 0.764 and better than Mordred’s 0.756. The results meet a predeclared pooled parity threshold but not all per‑assay thresholds, and further analysis confirms the competitiveness of the auditable representation while highlighting unresolved assay‑level differences.

By Yiqi Yao, Miquel Duran-Frigola
arXiv AI
Jul 21

Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion

arXiv:2607. 17601v1 Announce Type: cross Abstract: Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust.

By Yongchan Hong, Defu Cao, Wenjin Liu, Thomas Ku, Jordy Homing Lam, Emily Nguyen, Willie Neiswanger, Vsevolod Katritch, Yan Liu