Proteo-R1: Reasoning Foundation Models for De Novo Protein Design
arXiv:2605. 02937v2 Announce Type: replace-cross Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity.
arXiv:2605. 02937v2 Announce Type: replace-cross Abstract: Deep learning in de novo protein design has achieved atomic-level fidelity.
ChemVTS-Bench is a domain-authentic benchmark that evaluates Visual‑Textual‑Symbolic reasoning in multimodal large language models for chemistry. It presents diverse chemical problems—organic molecules, inorganic materials, and 3D crystal structures—in three input modes: visual-only, visual‑text hybrid, and SMILES-based symbolic. The benchmark includes an automated agent workflow for inference, answer verification, and failure diagnosis, and shows that visual-only inputs and structural chemistry remain challenging for current models.
arXiv:2411. 04440v1 Announce Type: cross Abstract: Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive.
arXiv:2607. 20557v1 Announce Type: cross Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition.
MolEmb is a lightweight framework that adapts multimodal large language models (MLLMs) to serve as general molecular embedding models. By aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective, MolEmb produces embeddings conditioned on both a molecular profile and a natural‑language semantic context. The model performs competitively on molecular property prediction and enables cross‑modal molecule‑text retrieval, while the newly introduced MolCAR benchmark demonstrates that context‑aware molecular embedding is largely a data property of the supervision.
arXiv:2605. 01625v3 Announce Type: replace Abstract: Proteins are inherently multiscale physical systems whose functional properties emerge from coordinated structural organization across multiple spatial resolutions, ranging from atomic interactions to global fold topology.
ChemMLLM is a unified chemical multimodal large language model designed for molecule understanding and generation across text, SMILES strings, and images. The authors curated five multimodal tasks and benchmarked ChemMLLM against leading general MLLMs, chemical LLMs, and specialized models, finding it outperforms general-purpose MLLMs and matches specialized models on all tasks. The study demonstrates that a single foundation model can handle diverse cross‑modal chemical tasks, including image generation, enabling more intuitive visual human‑AI interaction.
LatentVerse is a new framework that provides a web-based visual analytics platform and a command-line interface for analyzing multimodal latent representations. It unifies diagnostics for representation quality metrics and extends analysis to multimodal settings by decomposing embeddings into shared and modality-specific components. The authors evaluate the tool through simulations, real biomedical data analyses, and a user study, demonstrating its utility for interpretable evaluation of foundation model representations.
arXiv:2607. 07708v1 Announce Type: cross Abstract: Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization.
arXiv:2606. 03435v1 Announce Type: new Abstract: Cell Painting combines multiplexed fluorescent staining, high-content imaging, and quantitative analysis to generate high-dimensional phenotypic readouts to support diverse downstream tasks such as mechanism-of-action (MoA) inference, toxicity prediction, and construction of drug-disease atlases.
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order.
arXiv:2609.38879v1 Announce Type: cross Abstract: Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Prote...