arXiv AI

The Montparnasse Algorithm for RNA Design

arXiv:2606. 07562v1 Announce Type: cross Abstract: RNA design consists of discovering a nucleotide sequence that optimizes predefined criteria, such as secondary structure.

arXiv AI
Aug 19

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

The paper introduces MCTH (Monte Carlo Tree Hallucination), an inference-only framework that performs all‑atom biomolecular sequence‑structure co‑design by treating pretrained folding and inverse‑folding models as black‑box operators. MCTH uses Monte Carlo Tree Search to allocate a fixed inference budget across competing design trajectories, incorporating model confidence, uncertainty, and cross‑expert consensus. Experiments across protein‑RNA, protein‑DNA, protein‑protein, and protein‑ligand design show that adaptive search outperforms simpler sampling strategies, and evaluations with AlphaFold3 and Chai‑1 demonstrate transferability beyond the search‑time oracle.

By Xuefeng Liu, Mingxuan Cao, Xiao Luo, Songhao Jiang, Tobin Sosnick, Jinbo Xu, Louis Maher, Rick Stevens
arXiv Machine Learning
Sep 11

Sequence-Informed Geometric Evaluation of RNA 3D Structures

The paper introduces SIRGE, a sequence-informed geometric evaluator for RNA 3D structures that integrates nucleotide embeddings from a pretrained RNA language model into structural representations. SIRGE demonstrates superior performance over existing evaluators in Kendall–τ alignment, Top‑1 selection, and Top‑3 ranking. Controlled experiments reveal that sequence conditioning corrects errors of a purely geometric model and enhances target‑level ranking, suggesting that pretrained sequence representations provide complementary ranking information to geometric reasoning.

By Andrea Zerio, Yighua Yao, Alessandro Micheli, Roland G. Huber, Mile Sikic, Samir Bhatt, Andres R. Masegosa, Yuangang Pan
arXiv Machine Learning
Sep 7

Small Molecule Optimization with Large Language Models

The paper introduces Mol-E, an evolutionary algorithm that leverages large language models trained on molecular data to generate candidate molecules. Mol-E achieves state‑of‑the‑art performance on the Practical Molecular Optimization benchmark, scoring 17.500 in the task‑agnostic regime and 20.551 in the task‑informed regime. It also outperforms baseline methods in multi‑property optimization tasks involving docking against DRD2, MK2, and AChE.

By Philipp Guevorguian, Menua Bedrosian, Tigran Fahradyan, Gayane Chilingaryan, Armen Aghajanyan, Hrant Khachatrian
arXiv Machine Learning
Jul 17

A Machine Learning Benchmarking Framework for Lipid Nanoparticle Transfection Efficiency Prediction

arXiv:2507. 03209v2 Announce Type: replace-cross Abstract: The discovery of new ionizable lipids for efficient lipid nanoparticle (LNP)-mediated RNA delivery remains a major bottleneck in RNA therapeutics development.

By Asal Mehradfar, Mohammad Shahab Sepehri, Jose Miguel Hernandez-Lobato, Glen S. Kwon, Mahdi Soltanolkotabi, Salman Avestimehr, Morteza Rasoulianboroujeni
arXiv Machine Learning
Jun 30

Inference-time optimization for experiment-grounded protein ensemble generation

arXiv:2602. 24007v3 Announce Type: replace-cross Abstract: Protein function relies on dynamic conformational ensembles, yet current generative models like AlphaFold3 often fail to produce ensembles that match experimental data.

By Advaith Maddipatla, Anar Rzayev, Marco Pegoraro, Martin Pacesa, Paul Schanda, Ailie Marx, Sanketh Vedula, Alex M. Bronstein