arXiv AI By Kaihua Ding

Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data

Read the original on arXiv AI →

The paper investigates whether training classical machine learning models remains worthwhile when large language models (LLMs) can label tabular data without training. By defining a labeled‑data crossover point (N*) where a trained classical model surpasses a frozen LLM’s flat error, the authors analyze 126 student evaluations of GPT models across 18 datasets and compare them to power‑law learning curves of six classical model families. Results show that in 86% of cases a classical model outperforms the LLM with no more labeled data than already available, and the crossover occurs at a median of about 6% of the training set, suggesting that collecting a few hundred labels and training a gradient‑boosted model is typically advantageous.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction

arXiv:2606. 15314v1 Announce Type: cross Abstract: Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registered prototype will require a retrofit, which retrofit package it will need, and how long the work will take.

By Aina Vila Pons, Ioannis Tzachristas, Constantinos Antoniou