arXiv AI By Jorge L\'opez-Varela, J. Ignacio Hidalgo, Jos\'e-Manuel Mu\~noz, Omar Costilla-Reyes, Esther Maqueda, Jesus Moreno-Fernandez, Tom\'as Gonz\'alez-Vidal, J. Manuel Velasco, Oscar Garnica

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study

Read the original on arXiv AI →

The paper investigates using Large Language Models (LLMs) as post-hoc auditors for symbolic regression models produced by genetic programming. By having three LLMs analyze and rank four evolved expressions across multiple runs, the study compares the LLM-generated rankings and interpretations with assessments from three clinicians. Results show that LLMs provide more favorable comparative rankings than isolated term-level interpretations, yet they also generate physiologically and mathematically questionable explanations, suggesting they are best used under expert oversight rather than for autonomous validation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

Evolutional Math: Cross-Validated Island-Model Genetic Programming for Interpretable Symbolic Regression on Small, Wide Datasets

arXiv:2606. 28381v1 Announce Type: cross Abstract: Symbolic regression via genetic programming routinely fails on small, wide datasets - a regime common in clinical-trial monitoring, biostatistics, and engineering pilot studies - by converging on bloated, overfit expressions that exploit correlation rather than prediction.

By Artem Andrianov (Cyntegrity Germany GmbH, Hofheim am Taunus, Germany)
arXiv AI
Sep 7

A Semantic Model of Genetic Evidence: A Step Toward Bridging the Basic-Science-Clinic Gap

The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.

By Michael Bouzinier, Dmitry Etin
arXiv Machine Learning
Aug 27

InsightSR: Refining Symbolic Regression Search Spaces via Parallel Semantic and Structural LLM Guidance

InsightSR is a new framework that integrates Large Language Models (LLMs) with the PySR genetic programming engine to refine symbolic regression search spaces. It employs two LLM-guided pathways: a Semantic Seed Pathway that generates dimensionally consistent functional skeletons, and a Structural Feature Pathway that suggests nonlinear feature transformations. Over successive iterations, these pathways expand the input space and shift the search toward shallow, semantically informed trees, with a feedback loop that evaluates and refines candidate features. The method achieves a 95% exact recovery rate on the Feynman benchmark and 80.18% accuracy on the LLM-SRBench LSR-Transform task, outperforming existing genetic programming and neural-symbolic approaches while preserving strong out-of-distribution generalization.

By Yating Ling, Wenjing Cun, Zhitang Chen