arXiv Machine Learning

Explainable Machine Learning for Broadband Adoption Disparities: Tract-Level Prediction and SHAP-Based Factor Profiling

arXiv Machine Learning
Jul 21

STRATA: A Name-and-Geography Race Inference Model for Fair Lending and Housing Equity Applications

arXiv:2504. 21259v2 Announce Type: replace-cross Abstract: Accurate imputation of race and ethnicity (R&E) is essential for fair lending compliance under ECOA, HMDA, and the Community Reinvestment Act, where up to 15% of mortgage applications carry missing race data and regulated institutions bear responsibility for identifying disparities on those records.

By S. Chalavadi, A. Pastor, T. Leitch
arXiv Statistics ML
1d ago

Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States

The paper introduces DSP‑BART‑HS, a Dynamic Spatial Panel Bayesian Additive Regression Trees model with Horseshoe shrinkage, designed for high‑dimensional spatio‑temporal panel data. Across nine simulated scenarios, the model outperforms or matches a wide range of spatial econometric, non‑parametric machine learning, and small‑area estimators, especially when individual‑level non‑linearity drives outcome variance. The authors validate the method on two U.S. county‑level applications—intergenerational economic mobility and geographic income inequality—showing strong predictive accuracy even under unseen‑region, random, and temporal holdouts, while noting a temporal extrapolation advantage for a simpler autoregressive model.

By Hammed A. Olayinka, Saheed O. Olayemi
arXiv Machine Learning
Aug 11

Crowd-Sourced Geographies of Income: Using Google Maps Points of Interest as High-Frequency Proxies for Sub-Municipal Income Estimation in Sao Paulo, Brazil

arXiv:2608. 07871v1 Announce Type: cross Abstract: Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade.

By Adrienne C. Kinney, Anya Workman, Ademar Takeo Akabane, Jenna Barac, Paulo Fernando Braga Carvalho, Jeova Farias, Fernando Nascimento, Paulo Ricardo da Silva Oliveira
arXiv Machine Learning
Aug 18

In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults

arXiv:2608. 14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults.

By Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
arXiv Machine Learning
Aug 19

Predicting Male Domestic Violence Using Explainable Ensemble Learning and Exploratory Data Analysis

The paper presents a data‑driven study of male domestic violence (MDV) in Bangladesh, using exploratory data analysis to uncover patterns such as verbal abuse prevalence and the influence of financial dependency. It evaluates 10 traditional ML models, 3 deep learning models, and 2 ensemble models, ultimately proposing a stacking ensemble with ANN and CatBoost base classifiers and Logistic Regression meta‑model that achieves 95% accuracy and 99.29% AUC. Explainable AI techniques (SHAP, LIME) and statistical validation confirm the model’s superior performance and highlight key features driving predictions.

By Md Abrar Jahin, Saleh Akram Naife, Fatema Tuj Johora Lima, M. F. Mridha, Md. Jakir Hossen
arXiv AI
Aug 28

Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration

The paper presents a framework for integrating explainable AI into customer churn prediction for telecommunications. It benchmarks four classifiers—Logistic Regression, Random Forest, XGBoost, and LightGBM—on the IBM Telco Customer Churn dataset, finding comparable performance with Logistic Regression achieving the highest AUC-ROC and LightGBM the highest accuracy. Explanations are provided via SHAP and LIME at both global and instance levels, and a four‑layer CRM integration architecture is proposed to translate risk scores and attribution vectors into actionable retention strategies, projecting a 3.3–5.3 percentage point reduction in churn and $199K–$319K savings per campaign cycle.

By Sandeep Gaddamwar