arXiv Machine Learning By Bum Chul Kwon, Ben Shapira, Moshiko Raboh, Shreyans Sethi, Shruti Murarka, Joseph A Morrone, Leili Zhang, Wendy Cornell, Jianying Hu, Parthasarathy Suryanarayanan

STAR-VAE: A Scalable Latent-Variable Transformer for Controllable Molecular Generation

Read the original on arXiv Machine Learning →

STAR-VAE is a Transformer-based variational autoencoder that uses SELFIES encoding and a bidirectional encoder with an autoregressive decoder pretrained on 79 million PubChem molecules. It incorporates a property signal to jointly condition the prior, posterior, and decoder, and employs LoRA adapters for fine‑tuning on small datasets without altering the backbone. The model achieves 100 % validity and near‑perfect novelty in MOSES sampling, low KL divergence on several GuacaMol descriptors, strong synthetic‑accessibility conditioning, and effective docking‑score control across multiple protein targets, while also enabling scaffold recovery and diverse label‑conditioned generation on ChEMBL targets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

StabilityArc: Decoding Protein Sequence Embeddings into Generalizable Stability Landscapes

StabilityArc is a method that decodes protein sequence embeddings into generalizable stability landscapes. It uses a shared RoPE transformer to map frozen ESMC-600M residue representations into an Lx20 matrix of substitution effects, with a symmetric, contact-aware residual to predict epistasis. In extensive leave-one-protein-out tests on 134,794 ProteinGym variants, StabilityArc achieves a Spearman correlation of 0.7134, surpassing the best zero‑shot baseline, and further improves Kermut’s performance when used as a prior.

By Aaron L. Feller, Andrew D. Ellington, Claus O. Wilke
arXiv Machine Learning
Sep 4

SimpleDesign: A Joint Model for Protein Sequence and Structure Codesign

SimpleDesign is a single-stage, end-to-end model for joint protein sequence and structure design that eliminates the need for multi-stage training. It combines discrete cross-entropy for sequences with a regression objective for structures, using a Mixture-of-Transformer architecture to handle modality-specific processing while maintaining global self-attention. Trained on over 2 million sequence-structure pairs, SimpleDesign achieves strong performance on co-design and unconditional generation benchmarks.

By Jiarui Lu, Yuyang Wang, Yizhe Zhang, Jiatao Gu, Navdeep Jaitly, Joshua M. Susskind, Miguel \'Angel Bautista