arXiv Machine Learning
Aug 27

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

InsightSR is a new framework that integrates Large Language Models (LLMs) with the PySR genetic programming engine to refine symbolic regression search spaces. It employs two LLM-guided pathways: a Semantic Seed Pathway that generates dimensionally consistent functional skeletons, and a Structural Feature Pathway that suggests nonlinear feature transformations. Over successive iterations, these pathways expand the input space and shift the search toward shallow, semantically informed trees, with a feedback loop that evaluates and refines candidate features. The method achieves a 95% exact recovery rate on the Feynman benchmark and 80.18% accuracy on the LLM-SRBench LSR-Transform task, outperforming existing genetic programming and neural-symbolic approaches while preserving strong out-of-distribution generalization.

By Yating Ling, Wenjing Cun, Zhitang Chen
arXiv Machine Learning
Sep 21

MOSAIC-SR: Transformer-Guided Symbolic Regression for Scientific Equation Recovery

MOSAIC‑SR is a new symbolic regression method that combines a pretrained Transformer with search‑based refinement. The Transformer generates multiple initial sketches, which seed searches that jointly recover equation structure and constants using scale‑aware optimization and symbolic repair. On the SRSD‑Feynman dataset and six other benchmarks, MOSAIC‑SR achieves the highest symbolic solution rate and ranks among the top two in predictive accuracy, even when irrelevant dummy variables are present.

By Peiyi Zheng, Yanming Kang, Hans De Sterck, Giang Tran
arXiv Machine Learning
Aug 31

Probabilistic Symbolic Regression for Equation Discovery via Operator-induced and Regularized Symbolic Forests

The paper presents a probabilistic symbolic regression framework that models mathematical expressions as ensembles of symbolic trees, using a regularizing prior to control complexity and an Occam’s window-based posterior to capture uncertainty across plausible models. It provides theoretical guarantees on posterior concentration, including near‑parametric rates when an exact finite formula exists and oracle results under misspecification. Empirical results show the method outperforms state‑of‑the‑art competitors in predictive accuracy, symbolic complexity, and structural recovery on benchmark scientific equations and a materials discovery task.

By Somjit Roy, Pritam Dey, Bani K. Mallick, Debdeep Pati