arXiv AI

Auto-Formalizing Neuro-Symbolic Predictors

arXiv Machine Learning
Jun 16

Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification

arXiv:2601. 22642v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid.

By Chuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan, Zijian Zhao, Zhengyu Chen, Yuchen Tian, Lijun Wu, Conghui He, Sirui Han, Yike Guo
arXiv AI
Sep 4

RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data

RECAST is a new framework that generates datasets with far more constraints per example than existing benchmarks, aiming to push large language models (LLMs) to better follow complex instructions. The authors built RECAST-30K, a 30,000‑instance dataset covering 19 constraint types extracted from real prompt‑response pairs, and showed that fine‑tuning on it improves LLMs’ ability to handle complex tasks without harming general performance. RECAST also provides rule‑based and LLM‑based validators for automatic constraint verification, enabling reward‑based reinforcement learning to further enhance model performance on challenging tasks.

By Zhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu, Zisu Huang, Muzhao Tian, Jianhan Xu, Yuanzhe Shen, Qi Qian, Muling Wu, Xiaohua Wang, Changze Lv, He-Da Wang, Hu Yao, Xiaoqing Zheng, Xuanjing Huang
arXiv AI
Sep 25

Neuro-symbolic AI for Industrial Configuration

The paper "Neuro-symbolic AI for Industrial Configuration" discusses how Large Language Models (LLMs) fall short for industrial product configuration due to their probabilistic nature, which conflicts with the need for syntactically valid, semantically consistent outputs that align with extensive feature and rule knowledge bases. It proposes Neuro-symbolic (NeSy) AI as a promising solution, outlining three integration strategies—hybrid inference, hybrid fine‑tuning, and hybrid training—and presents a taxonomy of these approaches. The authors describe their efforts to implement a NeSy-based configuration copilot, derive practical design choices for trustworthy AI deployment in engineering settings, and highlight key research challenges, especially scaling NeSy methods from academic prototypes to full‑scale industrial configurators.

By Danilo Valerio, Philipp Kogler, Stefan Bischof, Thomas Hubauer, Huzefa Rangwala
arXiv Machine Learning
Sep 2

Neural Symbollic Regression Using Deep Learning and Sparse Modelling

Neural Symbolic Regression (NSR) uses neural networks as functional preconditioners to learn smooth, noise‑robust approximations of target functions in an interaction‑aware nonlinear feature space. A subsequent LASSO step extracts sparse, interpretable closed‑form expressions, while distributed hyperparameter optimization with Ray Tune and ASHA scheduling improves predictive accuracy and symbolic fidelity. Experiments on the Nguyen benchmark demonstrate that NSR outperforms SINDy and untuned neural baselines in RMSE, noise robustness, and out‑of‑distribution generalization, with ablation studies highlighting the importance of feature interactions, neural depth, and tuning strategies.

By Ravi Kumar U, Sumitra S
arXiv Machine Learning
Sep 11

Beyond Solver Verdicts: Generative Reward Models for Autoformalization

The paper identifies a new failure mode in neurosymbolic systems called Verdict‑Preserving‑Unfaithfulness (VPU), where incorrect formal encodings can still pass solver checks. It introduces Generative Verification (GenV), a method that uses a language model to produce a continuous reference‑equivalence score without relying on explicit localization. Experiments show GenV+HN achieves high AUROC, generalizes to unseen translators, and improves downstream agent performance by 11.3 points.

By Vikash Singh, Debargha Ganguly, Aman Goel, Ali Torkamani, Xiaoxue Han, Joseph Lilien, Ferhat Erata, Vipin Chaudhary
arXiv AI
Sep 17

An Agentic Framework for Neuro-Symbolic Programming

The paper introduces AgenticDomiKnowS (ADS), a framework that converts free‑form task descriptions into fully functional DomiKnowS neuro‑symbolic programs. ADS employs an agentic workflow that builds and tests each DomiKnowS component independently, optionally allowing human‑in‑the‑loop refinement. The authors demonstrate that ADS enables both experienced and novice users to create complete neuro‑symbolic programs in 10–15 minutes, compared to the hour required for manual coding.

By Aliakbar Nafar, Chetan Chigurupati, Danial Kamali, Hamid Karimian, Parisa Kordjamshidi