arXiv:2607. 23286v1 Announce Type: new Abstract: Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective is to discover predictive feature transformations from an exponentially large search space.
By Sha Li, Naren Ramakrishnan
arXiv:2606. 02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures.
By Andrej Tschalzev, Nick Erickson, Yuyang Wang, Huzefa Rangwala, Stefan L\"udtke, Heiner Stuckenschmidt, Christian Bartelt
arXiv:2607. 01548v1 Announce Type: cross Abstract: Large language models are increasingly used as open-ended search operators in evolutionary optimization.
By Ege Onur Taga, Yilin Zhuang, M. Emrullah Ildiz, Petros Mol, Abhimanyu Das, Karthik Duraisamy, Samet Oymak
The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.
By Merwan Barlier, Blaz Skrlj
Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature Engineering (EFE), a framework for using LLM-based evolution to discover preprocessing transformations for structured data.
SymboLLM-FE combines symbolic regression and large language models to automate feature engineering for tabular data. It first extracts mathematically expressive formulas that correlate strongly with the target, then refines them with LLMs to improve interpretability. Experiments on six real‑world datasets and four Kaggle competitions show that SymboLLM‑FE outperforms existing AutoFE methods while reducing the number of costly LLM calls.
By Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo