The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.
By Md Rezwanul Islam, Wael Mohammed
The paper reports a previously undocumented failure of global gradient‑boosted tree forecasters when applied to hierarchical aggregates. Training a single tree on individual series causes the model to predict a constant outside its training range, leading to severe under‑prediction of the total (30–50× in production and up to 496× in a public M5 reconstruction). The authors characterize this collapse across five datasets, three tree libraries, and multiple training seeds, and demonstrate that simple preprocessing steps—per‑series scaling, weighted aggregate‑level training, or seasonal differencing—can prevent it.
By Md Rezwanul Islam, Wael Mohammed
The paper audits the IBM Telco Customer Churn benchmark, revealing that common practices inflate performance metrics. It shows that pre‑split SMOTE boosts churn‑class F1 by 13.1 points, that isotonic regression is the best calibration method while temperature scaling fails on tree ensembles, and that the cost‑optimal decision threshold is 5–10 times lower than the F1‑optimal one, saving about $77,000 per 1,000 customers. The authors also test generalisation on Iranian Telecom and Bank churn datasets, and propose a four‑component reporting checklist with reproducible code.
By Soumyadeep Roy
The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.
By Sandeep Gaddamwar
arXiv:2608. 14903v1 Announce Type: new Abstract: Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, release time, or expert belief.
By Fabricio F Costa
The paper reports that language‑model agents used for customer‑relationship management can be misled by optimistic assertions from sales representatives in CRM records, leading to incorrect deal approvals. In a benchmark of 100 lead‑qualification tasks, models incorrectly cleared 29 of 31 deals where the representative’s claims contradicted company policies, with misalignment rates ranging from 87% to 97% across seven models. The authors propose diagnostic methods—including bucket analysis, same‑information controls, and compute‑step controls—to distinguish persuasion from information gaps and to quantify the impact of incentive‑misaligned witnesses.
By Rahul Balakavi