arXiv Machine Learning

Predictive Query Language: A Domain-Specific Language for Predictive Modeling on Relational Databases

arXiv:2602. 09572v3 Announce Type: replace-cross Abstract: The purpose of predictive modeling on relational data is to predict future or missing values in a relational database, for example, future purchases of a user, risk of readmission of the patient, or the likelihood that a financial transaction is fraudulent.

arXiv Machine Learning
Jul 16

Foundation Models for Credit Risk Prediction: A Game Changer?

arXiv:2605. 18147v2 Announce Type: replace Abstract: Predictive models play a pivotal role in credit risk management, guiding critical decisions through accurate estimation of default probabilities and losses.

By Bart Baesens, Andreas Goethals, Stefan Lessmann, Simon De Vos, Cristi\'an Bravo, David Martens, Victor Medina-Olivares, Christophe Mues, Maria Oskarsd\'ottir, Seppe vanden Broucke, Tony Van Gestel, Tim Verdonck, Wouter Verbeke
arXiv Computation and Language
Sep 11

LLMAR: A Tuning-Free Recommendation Framework for Sparse and Text-Rich Industrial Domains

LLMAR is a tuning‑free recommendation framework designed for sparse, text‑rich industrial B2B domains. It transforms user behavioral history into structured semantic motives using LLM inference, employs a reflection loop to self‑correct hallucinations, and operates cost‑effectively with asynchronous batch processing. Experiments on MovieLens‑1M, Amazon Prime Pantry, and a construction risk dataset show LLMAR surpasses state‑of‑the‑art learning models, achieving up to a 54.6% nDCG@10 improvement while keeping inference costs around $1 per 1,000 users.

By Ryogo Hishikawa, Ichiro Kataoka, Shinya Yuda
arXiv AI
Sep 15

Towards Optimizing SQL Generation via LLM Routing

The paper "Towards Optimizing SQL Generation via LLM Routing" proposes a routing approach for Text-to-SQL tasks that dynamically selects the most cost‑effective large language model (LLM) for each query. Two routing strategies—score‑based and classification‑based—are introduced, achieving accuracy comparable to the best LLM while reducing latency and monetary cost. The authors design the routers for easy training and efficient inference, and demonstrate a practical accuracy‑cost trade‑off on the BIRD dataset.

By Mohammadhossein Malekpour, Nour Shaheen, Foutse Khomh, Amine Mhedhbi
arXiv Machine Learning
1d ago

STEER: Reducing Inference Cost in Relational Foundation Models through Semantically Informed Sampling

STEER is a sampling method for relational foundation models that reduces inference cost by focusing on the most relevant tables for a prediction task. It uses a large language model to rank foreign‑key edges in the database schema into relevance tiers, then assigns traversal probabilities based on these tiers. Evaluated on three state‑of‑the‑art RFMs, STEER cuts inference context size by roughly 40% on average while preserving or improving accuracy.

By Abdalla Mohamed, Ashraf Aboulnaga