arXiv AI

Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling

arXiv:2607. 18515v1 Announce Type: cross Abstract: This study contributes toward development of an Automated Data Processing (ADP) framework designed to evaluate and reinforce optimal machine learning model-feature combinations for predictive tasks in fused deposition modeling (FDM) process datasets.

arXiv Machine Learning
Aug 31

SymboLLM-FE: LLM-Accelerated Symbolic Regression for Automated Feature Engineering on Tabular Data

SymboLLM-FE combines symbolic regression and large language models to automate feature engineering for tabular data. It first extracts mathematically expressive formulas that correlate strongly with the target, then refines them with LLMs to improve interpretability. Experiments on six real‑world datasets and four Kaggle competitions show that SymboLLM‑FE outperforms existing AutoFE methods while reducing the number of costly LLM calls.

By Zi-Jian Cheng, Zi-Yi Jia, Zhi Zhou, Yu-Feng Li, Lan-Zhe Guo
arXiv AI
Jun 17

TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins

arXiv:2606. 17660v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and na\"ive runs can even degrade model performance.

By Yuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan, Xiaotian Lin, Yanwei Xu, Yuyu Luo, Weikai Yang, Nan Tang
arXiv Machine Learning
4d ago

Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients

The paper introduces Non-Myopic Active Feature Acquisition via Pathwise Policy Gradients (NM-PPG), a method that relaxes the feature acquisition process to allow continuous, low‑variance policy gradients over the entire acquisition trajectory. It incorporates a straight‑through rollout that mimics discrete acquisitions during inference while enabling end‑to‑end training, and provides an average‑case upper bound on gradient variance to guide temperature sharpening. Experiments on synthetic and real datasets show that NM-PPG outperforms existing active feature acquisition baselines.

By Linus Aronsson, Morteza Haghir Chehreghani
arXiv AI
Aug 28

Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study

The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.

By Chorok Lee