arXiv AI

PhyMo: A Physical-Field Modality for Multimodal AI4Physics

The paper introduces PhyMo, a physics‑grounded multimodal framework that uses a physical‑field modality to represent heterogeneous measurements via PDE‑associated operators. It follows a three‑stage learning process: pretraining a physical‑field encoder with PDE residual supervision, aligning its representations with visual embeddings in a shared latent space, and applying downstream prediction heads to the fused multimodal representations. Experiments on five diverse physical datasets show that PhyMo outperforms the strongest baseline on each dataset, establishing its effectiveness for multimodal representation learning in AI for Physics.

arXiv AI
Jul 29

COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations

arXiv:2501. 17296v4 Announce Type: replace-cross Abstract: Multiphysics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domains Although neural operators especially the Fourier Neural Operator FNO have significantly improved computational efficiency they often fail to effectively capture intricate correlations inherent in coupled physical processes To address this limitation we introduce COMPOL a novel coupled multiphysics operator learning framework COMPOL extends conventional operator architectures by incorporating sophisticated recurrent and attentionbased aggregation mechanisms effectively modeling interdependencies among interacting physical processes within latent feature spaces Our approach is architectureagnostic and seamlessly integrates into various neural operator frameworks that involve latent space transformations Extensive experiments on diverse benchmarksincluding biological reactiondiffusion systems patternforming chemical reactions multiphase geological flows and thermohydromechanical processes demonstrate that COMPOL consistently achieves superior predictive accuracy compared to stateoftheart methods.

By Junqi Qu, Tao Wang, Yushun Dong, Hewei Tang, Shibo Li
arXiv AI
Jul 2

A Multi-Resolution Finite-Volume Inspired Deep Learning Framework for Spatiotemporal Dynamics Prediction

arXiv:2607. 00460v1 Announce Type: cross Abstract: Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven neural networks that suffer from high training costs, error accumulation, and limited generalizability to unseen parameters.

By Xin-Yang Liu, Xiantao Fan, Jian-Xun Wang
arXiv Computation and Language
Aug 27

OmniPhys: A Unified Multimodal Benchmark for Physics Understanding and Generation from Chinese Educational Corpora

OmniPhys is a large-scale multimodal benchmark designed to evaluate physics understanding and reasoning in models. It contains 15,246 questions and 19,850 images sourced from Chinese educational materials ranging from middle school to university level, with detailed annotations for fine-grained analysis. The benchmark also tests models’ ability to generate structured physics diagrams, a key component of authentic problem solving, and highlights gaps in current multimodal large language models.

By Hao Chen, Yumin Lin, Nadila Yushanjiang, Xin Lin, Min Zhang
arXiv AI
Jul 24

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

arXiv:2607. 20557v1 Announce Type: cross Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition.

By Hesen Chen, Xinyu Su, Xiaomeng Yang, Yuetan Lin, Zixiong Yang, Junyi An, Fenglei Cao, Yifeng Jiao, Yunqi Zhang, Yuan Cheng, Zhiyu Tan, Hao Li, Libo Wu, Yuan Qi
arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
arXiv Machine Learning
Sep 16

Can Deep Learning Achieve Cross-Physics Mapping?

The paper introduces Cross-Physics Mapping (CPM), an operator-learning framework that enables deep learning to translate physical fields governed by different equations. By aligning latent representations and applying a dimensionless scaling principle, CPM maps between heterogeneous domains such as diffusion and wave fields. Experiments with seven neural operator architectures show directional asymmetry: diffusion-to-wave mapping is harder, while wave-to-diffusion mapping is more stable, with neural operators outperforming conventional convolutional baselines.

By Pengfei Zhu, Julien Lecompagnon, Mathias Ziegler