arXiv AI

Toward autocorrection of chemical process flowsheets using large language models

arXiv:2312. 02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process flows and equipment configurations.

arXiv Machine Learning
Sep 14

SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling

SAGE-Loop is a new closed‑loop, self‑adaptive AutoML framework that uses large language models to generate and validate machine learning pipelines in multiple rounds, allowing trial‑and‑repair and adaptive ensemble selection for both supervised and unsupervised tasks. It addresses the lack of instant feedback and correction in existing AutoML by enabling process‑level recovery from failures and dynamic use of model diversity. Experiments on 20 public datasets show consistent improvements in performance and stability across classification, regression, and clustering, and demonstrate the system’s ability to recover from execution failures.

By Junquan Gu, Shibo Cui, Xiangfeng Luo, Hang Yu
arXiv AI
Sep 24

PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

PotARCin expands the ARC benchmark by evaluating abstract reasoning across five dimensions—Definition, Classification, Constrained Generation, Editing, and Inversion—using programmatic generation of new task instances. The study shows a 25‑52 percentage‑point performance gap between standard ARC evaluation and PotARCin, and reveals that multi‑dimensional assessment can reorder models that appear equivalent under single‑metric accuracy. Additionally, a new held‑out set, P‑ARC, demonstrates low model accuracy (1‑8%) across all dimensions, highlighting the need for more comprehensive tests of abstract reasoning.

By Claas Beger, Ryan Yi, Melanie Mitchell
arXiv AI
Jun 3

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

arXiv:2606. 02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datasets have never been rigorously audited.

By Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani, Angelo Montanari, Nicola Saccomanno
arXiv AI
Jul 15

A Neurosymbolic Approach to Natural Language Formalization and Verification

arXiv:2511. 09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and health-care that operate under strict policies.

By Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, R\'emi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovi\'c, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joseph Lilien, Divya Raghunathan, Zvonimir Rakamari\'c, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao
arXiv Machine Learning
Aug 27

iFlip: Iterative Feedback-driven Counterfactual Example Refinement

iFlip is an iterative refinement method for generating counterfactual examples using large language models. It incorporates three feedback types—model confidence, feature attribution, and natural language—to guide successive edits. Experiments show iFlip outperforms five state‑of‑the‑art baselines, achieving a 57.8% higher validity rate and improving model performance through counterfactual data augmentation.

By Yilong Wang, Qianli Wang, Nils Feldhus
arXiv AI
Sep 10

HoarePrompt: Structural Reasoning About Program Correctness in Natural Language

HoarePrompt is a new method that applies program verification concepts to natural language requirements, using large language models to generate step‑by‑step natural language descriptions of program states. It incorporates a few‑shot k‑induction technique to handle loops and then evaluates whether the annotated program satisfies the requirements. On the CoCoClaNeL dataset, HoarePrompt raises the Matthews correlation coefficient by 61% over zero‑shot chain‑of‑thought prompts and by 106% over test‑generation classifiers, with the inductive reasoning component adding a 26% MCC improvement.

By Dimitrios Stamatios Bouras, Yihan Dai, Tairan Wang, Yingfei Xiong, Sergey Mechtaev