When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization
arXiv:2607. 18573v1 Announce Type: new Abstract: Delay-risk models are usually judged by predictive accuracy.
Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones should a manager check first?
arXiv:2607. 18573v1 Announce Type: new Abstract: Delay-risk models are usually judged by predictive accuracy.
arXiv:2609.13840v1 Announce Type: new Abstract: A contract-logistics spare-parts operator is paid on order-level service: an order counts only if every requested line is fulfilled, yet forecasters ar...
The paper introduces SHAP concentration as a pre‑deployment diagnostic for detecting when conformal prediction will fail under distribution shift, specifically in gradient‑boosted classifiers. Using a COVID‑19 supply‑chain case study, the authors show that higher feature‑importance concentration correlates with larger drops in coverage, while standard shift detectors cannot differentiate between catastrophic and robust outcomes. The diagnostic is validated on additional datasets, and a formal theorem links concentration to worsening conformity‑score bounds, though it does not capture global‑sensitivity failures in neural networks.
The paper audits the IBM Telco Customer Churn benchmark, revealing that common practices inflate performance metrics. It shows that pre‑split SMOTE boosts churn‑class F1 by 13.1 points, that isotonic regression is the best calibration method while temperature scaling fails on tree ensembles, and that the cost‑optimal decision threshold is 5–10 times lower than the F1‑optimal one, saving about $77,000 per 1,000 customers. The authors also test generalisation on Iranian Telecom and Bank churn datasets, and propose a four‑component reporting checklist with reproducible code.
arXiv:2608.27704v1 Announce Type: new Abstract: When machine learning classifiers are retrained, inputs correctly classified by the previous model version may be misclassified by the updated version,...
arXiv:2607. 18530v1 Announce Type: cross Abstract: Supplier lead time forecasting is a central input to material requirements planning, inventory optimization, and supply chain risk management.
arXiv:2512. 18390v2 Announce Type: replace Abstract: Organizations often have an incumbent predictive model in production when new data sources become available.
arXiv:2607. 20655v1 Announce Type: cross Abstract: Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often underperform in production.
arXiv:2602. 06136v2 Announce Type: replace Abstract: Test-time adaptation (TTA) offers a compelling remedy for machine learning (ML) models that degrade under domain shifts, improving generalisation on-the-fly with only unlabelled samples.
arXiv:2608. 11154v1 Announce Type: new Abstract: Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value.
arXiv:2606. 03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment.
The paper introduces a regime‑diagnosis framework for industrial time‑series forecasting, highlighting that canonical loss functions embed fixed statistical priors that are violated in real‑world demand regimes such as zero‑inflation, skewness, and high variability. It proposes the Regime‑wise Relative Bias Vector (RBV) as a metric‑agnostic diagnostic that decomposes bias into an intrinsic floor and an excess attributable to training. A large‑scale study across 13 loss objectives and 60,000+ series demonstrates that regime‑aware diagnosis distinguishes optimization‑from‑bias failures and that regime‑aware training can eliminate pooling‑induced bias that mere capacity scaling cannot.