arXiv AI By Sirui Lu, Erickson Tjoa, J. Ignacio Cirac

Multi-agent Autoformalization of Tensor Network Theory

Read the original on arXiv AI →

arXiv:2607. 07857v1 Announce Type: cross Abstract: We build a team of specialized large language-model agents and present an agent-driven workflow for research-level formalization in theoretical physics, with the autoformalization of the fundamental theorem of matrix-product states as a demonstration.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

Long-horizon autoformalization of a core theorem underlying MIP* = RE

FormalFlow is a system that coordinates AI proving agents under human supervision to tackle long‑horizon formalizations, using a shared blueprint for nested planning, proving, and review loops. The team used it to produce a machine‑checked Lean 4 proof of the quantum soundness of the classical low‑individual‑degree test, a core theorem underlying MIP* = RE, in 63 days. The resulting library contains 126,367 lines of Lean code, all generated by agents, and corrects side conditions while preserving the published error bound under corrected assumptions.

By Sirui Lu, Ruixuan Deng, Yanqiao Zhu, Zhengfeng Ji
arXiv AI
Sep 7

AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics

AxQM is a new benchmark for formal proof synthesis in physics, comprising 1,019 Lean‑kernel‑checkable tasks drawn from the textbook *Quantum Computation and Quantum Information* by Nielsen and Chuang. The tasks are defined in a custom Lean library for finite‑dimensional quantum mechanics and are guaranteed solvable because they stem from a near‑complete formalization of the textbook’s formal content. Grading is deterministic, requiring proofs to compile, contain no sorrys, and introduce no new axioms.

By Weichen Winston Yin, Jacob M. Taylor, Dirk R. Englund, Frank H. L. Koppens
arXiv AI
Jun 9

Advancing Mathematics Research with AI-Driven Formal Proof Search

arXiv:2605. 22763v2 Announce Type: replace Abstract: Large language models (LLMs) increasingly excel at mathematical reasoning, but their unreliability limits their utility in mathematics research.

By George Tsoukalas, Anton Kovsharov, Sergey Shirobokov, Anja Surina, Moritz Firsching, Gergely B\'erczi, Francisco J. R. Ruiz, Arun Suggala, Adam Zsolt Wagner, Eric Wieser, Lei Yu, Aja Huang, Mikl\'os Z. Horv\'ath, Andrew Ferraiuolo, Henryk Michalewski, Edward Lockhart, Codrut Grosu, Thomas Hubert, Matej Balog, Pushmeet Kohli, Swarat Chaudhuri