arXiv Machine Learning

Zero-shot generalization of transformer neural operators to larger domains

arXiv:2606. 14597v1 Announce Type: new Abstract: Transformer-based neural operators have shown remarkable performance for approximating solution operators of partial differential equations on complex geometries.

arXiv AI
Aug 17

ArGEnT: Arbitrary Geometry-encoded Transformer for Operator Learning

arXiv:2602. 11626v3 Announce Type: replace-cross Abstract: Learning solution operators on arbitrary geometries remains a central challenge in scientific machine learning, especially for many-query simulation, physics-informed learning, and evolving geometries requiring accurate, geometry-aware predictions at arbitrary spatial locations.

By Wenqian Chen, Zhi-Feng Wei, Yucheng Fu, Michael Penwarden, Pratanu Roy, Panos Stinis
arXiv Machine Learning
Sep 17

HiLNO: A Hierarchical Latent Neural Operator with Multi-Scale Supervision for PDEs on General Geometries

HiLNO is a hierarchical latent neural operator that builds a fine‑to‑coarse‑to‑fine latent space and incorporates multi‑scale supervision and anisotropic Gaussian attention to preserve spatial information in PDE solutions with multiscale structures. The hierarchy reduces information loss during compression, while multi‑scale supervision aligns intermediate predictions with downsampled targets, and anisotropic attention facilitates feature transfer across scales. Experiments on standard PDE benchmarks and a large‑scale automotive aerodynamics task show that HiLNO achieves competitive accuracy while cutting parameter count by 84.4% and FLOPs by 69.2% compared with LinearNO, and it generalizes effectively to unseen spatial resolutions.

By Zhicheng Hu, Jiacheng Li, Min Yang
arXiv Computer Vision
Aug 31

Video Generative Models as Geometry Learner

The paper introduces GeoNeXt, a framework that repurposes pretrained video generative models for geometry estimation by framing it as a next‑frame prediction task. Unlike prior methods that either train separate depth/normal models or fine‑tune image diffusion backbones, GeoNeXt jointly models images and geometric targets, leveraging the structured knowledge of video models for more data‑efficient learning. Experiments show zero‑shot monocular depth and surface normal estimation that outperforms existing generative approaches and rivals discriminative state‑of‑the‑art methods while using far less training data.

By Haosen Yang, Jifei Song, Zhensong Zhang, Xiatian Zhu, Jiankang Deng