arXiv Machine Learning By Fan Yang, Matt Thomson

What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery

Read the original on arXiv Machine Learning →

The paper proposes a new data interpretation stage that transforms spatiotemporal field data into physically meaningful quantities before feeding them to a large language model for partial differential equation (PDE) discovery. On simulated benchmarks, this approach nearly triples the accuracy of recovered equations compared to using raw data, while incurring negligible computational cost and requiring no additional training. The method enables language models to read field data as a theorist would, facilitating automated field‑theory construction that can evolve alongside experimental data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 25

Decoupled Physical Modeling and Execution for Physics Reasoning

The paper introduces a framework that separates physical modeling from execution in physics reasoning tasks. It uses a two‑stage post‑training approach: supervised fine‑tuning to build structured models and reinforcement learning with rubric‑based feedback to refine them. Experiments on PhysReason, PhyX, and SeePhys show that this explicit modeling improves reasoning performance by about 3% on average for small LLMs.

By Ye Zhang, Xuehang Guo, Rui Pan, Pengfei Yu, Denghui Zhang, Manling Li, Qingyun Wang
Hugging Face Trending Papers
Jul 8

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order.

arXiv AI
Sep 2

MMAI Gym for Science: Training Liquid Foundation Models for Drug Discovery

arXiv:2603.03517v2 Announce Type: replace-cross Abstract: General-purpose large language models (LLMs) that rely on in-context learning do not reliably deliver the scientific understanding and perfor...

By Maksim Kuznetsov, Zulfat Miftahutdinov, Rim Shayakhmetov, Mikolaj Mizera, Roman Schutski, Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Thomas MacDougall, Mathieu Reymond, Mihir Bafna, Kaeli Kaymak-Loveless, Eugene Babin, Maxim Malkov, Mathias Lechner, Ramin Hasani, Alexander Amini, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
arXiv AI
Jul 9

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

arXiv:2607. 07708v1 Announce Type: cross Abstract: Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization.

By Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai
arXiv AI
Sep 4

Discovering High Level Patterns from Simulation Traces

The paper proposes an unsupervised learning approach that uses program synthesis to translate detailed simulation traces into sparse, high‑level structural patterns, making them easier for large language models to interpret. These pattern detectors can be guided by human‑provided labels such as "rigid collision" or "stretching spring" and produce transparent, explainable functions mapping system states to concise annotations. Experiments on a physics benchmark show that the annotated representations improve natural language reasoning about specific physical systems and enable natural‑language goals to be converted into reward programs for solution search.

By Sean Memery, Kartic Subr