arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.
By S. Chalavadi, A. Pastor, T. Leitch
The paper introduces DSP‑BART‑HS, a Dynamic Spatial Panel Bayesian Additive Regression Trees model with Horseshoe shrinkage, designed for high‑dimensional spatio‑temporal panel data. Across nine simulated scenarios, the model outperforms or matches a wide range of spatial econometric, non‑parametric machine learning, and small‑area estimators, especially when individual‑level non‑linearity drives outcome variance. The authors validate the method on two U.S. county‑level applications—intergenerational economic mobility and geographic income inequality—showing strong predictive accuracy even under unseen‑region, random, and temporal holdouts, while noting a temporal extrapolation advantage for a simpler autoregressive model.
By Hammed A. Olayinka, Saheed O. Olayemi
arXiv:2608. 07871v1 Announce Type: cross Abstract: Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade.
By Adrienne C. Kinney, Anya Workman, Ademar Takeo Akabane, Jenna Barac, Paulo Fernando Braga Carvalho, Jeova Farias, Fernando Nascimento, Paulo Ricardo da Silva Oliveira
arXiv:2608. 14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults.
By Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
arXiv:2609.27654v1 Announce Type: cross
Abstract: Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending...
By Sultan Amed, Tanmay Sen, Sayantan Banerjee
arXiv:2607. 07060v1 Announce Type: cross Abstract: Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly.
By Srikumar Krishnamoorthy
arXiv:2510. 22266v3 Announce Type: replace-cross Abstract: Identifying the factors that influence student performance in basic education is a central challenge for formulating effective public policies in Brazil.
By Rodrigo Tertulino, La\'ercio Alencar
The paper presents a data‑driven study of male domestic violence (MDV) in Bangladesh, using exploratory data analysis to uncover patterns such as verbal abuse prevalence and the influence of financial dependency. It evaluates 10 traditional ML models, 3 deep learning models, and 2 ensemble models, ultimately proposing a stacking ensemble with ANN and CatBoost base classifiers and Logistic Regression meta‑model that achieves 95% accuracy and 99.29% AUC. Explainable AI techniques (SHAP, LIME) and statistical validation confirm the model’s superior performance and highlight key features driving predictions.
By Md Abrar Jahin, Saleh Akram Naife, Fatema Tuj Johora Lima, M. F. Mridha, Md. Jakir Hossen
arXiv:2607. 15446v1 Announce Type: new Abstract: The cost of healthcare remains a concern in the United States and may have been influenced by disruptions associated with the COVID-19 pandemic.
By Alexey Kresin, Zien Cheng, Ammar Ahad, Ebiyomare Kelvin, Manish Sivaratri, Prabhjeet Singh, Omar Aljawfi, Olabisi Ojo, Nawar Shara
The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.
By Sandeep Gaddamwar
arXiv:2608. 08202v1 Announce Type: new Abstract: Data-centric curation pipelines frequently rely on model confidence scores to flag and filter noisy or mislabeled training instances.
By Sai Srikar Boddupalli
arXiv:2608.21567v1 Announce Type: new
Abstract: Human mobility serves as an essential proxy for understanding social, economic, and environmental dynamics in urban systems. Geospatial transferability...
By Zhiyong Zhou, Song Gao, Qianheng Zhang, Feng Zhang, Zhenhong Du