arXiv:2606. 04165v1 Announce Type: cross Abstract: High-precision calorimeter simulation at current and future colliders imposes rapidly growing computational demands, motivating the development of machine-learning surrogates for traditional Monte Carlo tools such as Geant4.
By Cheng Jiang, Sitian Qian, Kevin Pedro, Oz Amram, Huilin Qu, Maggie Voetberg
arXiv:2601. 21026v2 Announce Type: replace-cross Abstract: Sampling configurations at thermodynamic equilibrium is a central challenge in statistical physics.
By Louis Grenioux, Maxence Noble
arXiv:2505. 22391v2 Announce Type: replace-cross Abstract: Modeling physical systems in a generative manner offers several advantages, including the ability to handle partial observations, generate diverse solutions, and address both forward and inverse problems.
By Yi Zhang, Peng Wang, Difan Zou
arXiv:2601. 21284v2 Announce Type: replace-cross Abstract: Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature limits applicability in engineering and scientific problems where physical laws must be respected.
By Tianyi Zeng, Tianyi Wang, Jiaru Zhang, Zimo Zeng, Feiyang Zhang, Yiming Xu, Sikai Chen, Junfeng Jiao, Christian Claudel, Xinbo Chen
arXiv:2601.11716v2 Announce Type: replace-cross
Abstract: Accurate and efficient detector simulation is essential for modern collider experiments. To reduce the high computational cost, various fast...
By Thorsten Buss, Henry Day-Hall, Frank Gaede, Gregor Kasieczka, Katja Kr\"uger
arXiv:2608. 06107v1 Announce Type: new Abstract: Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fast, differentiable surrogate models.
By Guillaume Couairon, Alexis Jacq, Yu-Han Wu, Renu Singh, Yana Hasson, Quentin Berthet, Romuald Elie
The paper introduces the first controlled benchmark for optimizers in discrete diffusion models, evaluating seven optimizers (AdamW, Lion, Muon, SOAP, MARS, MARS‑M, Schedule‑Free) across four diffusion formulations: masked diffusion on text8, uniform diffusion on QM9 and LM1B, and Gaussian diffusion on CelebA‑64. Each optimizer undergoes the same search protocol and is retrained with full budget and multiple seeds, revealing that AdamW, while strong, is not universally optimal and that optimizers validated on autoregressive language models (Muon, MARS‑M, SOAP) can outperform tuned AdamW on certain tasks.
By Arman Bolatov, Egor Shulgin, David Li, Abduragim Shtanchaev, Sebastian U. Stich, Maxim Panov, Eric Moulines, Peter Richt\'arik, Martin Tak\'a\v{c}
arXiv:2605. 08318v2 Announce Type: replace Abstract: We study the problem of \emph{architecture selection} for deep learning models trained to solve partial differential equations (PDEs), asking when transformer-based architectures with learned attention outperform Fourier-domain neural operators.
By Brandon Yee, Pairie Koh, Jack Rodriguez, Mihir Tekal
Transolver‑σ is a neural PDE solver that jointly models spectral and physical subspaces to improve accuracy in both one‑step and autoregressive rollouts. The method uses adaptive physical-state interactions, Slice‑Residual Physics‑Attention, and an axis‑factorized Fourier operator to enable information exchange between representations. Across five standard PDE benchmarks, Transolver‑σ reduces benchmark‑averaged relative error by 33.4% compared to the strongest baseline and shows strong performance on coupled multiphysics systems and real‑world fluid and combustion data.
By Haonan Shangguan, Hang Zhou, Haixu Wu, Yuezhou Ma, Jianmin Wang, Mingsheng Long
arXiv:2609.30988v1 Announce Type: new
Abstract: Real-world image super-resolution (SR) requires recovering perceptually realistic high-resolution images from complex low-resolution observations while...
By Xin Di, Mingyu Shi, Yuanfei Bao, Long Peng, Yue Zhao, Jiaming Guo, Renjing Pei, Xueyang Fu, Yang Cao, Zheng-Jun Zha
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2607. 13682v2 Announce Type: cross Abstract: Radiative Gaussian splatting reconstructs sparse-view CT fast and accurately, and recent work attaches per-Gaussian posteriors to yield per-voxel uncertainty maps.
By Chulin Zhao, Yiran Xu, Shu Liu