arXiv Statistics ML

Dynamic Spatial Bayesian Machine Learning Model: Applications to Intergenerational Economic Mobility and Geographic Income Inequality in the United States

The paper introduces DSP‑BART‑HS, a Dynamic Spatial Panel Bayesian Additive Regression Trees model with Horseshoe shrinkage, designed for high‑dimensional spatio‑temporal panel data. Across nine simulated scenarios, the model outperforms or matches a wide range of spatial econometric, non‑parametric machine learning, and small‑area estimators, especially when individual‑level non‑linearity drives outcome variance. The authors validate the method on two U.S. county‑level applications—intergenerational economic mobility and geographic income inequality—showing strong predictive accuracy even under unseen‑region, random, and temporal holdouts, while noting a temporal extrapolation advantage for a simpler autoregressive model.

arXiv Machine Learning
Sep 10

Semi-Supervised Learning under Spatially Biased Sampling

arXiv:2609.07982v1 Announce Type: new Abstract: Standard semi-supervised learning (SSL) typically relies on labelled and unlabelled data sharing a common marginal distribution. This assumption is oft...

By Bright Wiredu Nuakoh, Francky Fouedjio, Stephen Bradshaw, Yaw Kwaafo Awuah-Mensah, Wei Hong Tan, Emet Arya, Ebenezer Afrifa-Yamoah
arXiv Machine Learning
Sep 2

Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment

The study investigates whether large language models (LLMs) can predict neighborhood-level human mobility without training data. Using anonymized Cuebiq data across four U.S. metropolitan areas, the authors compare zero‑shot LLM predictions to supervised baselines for various mobility outcomes and assess structural alignment with empirical trends. Results show supervised models outperform LLMs (average accuracy 0.580 vs. 0.435), with LLMs relying on coarse, stable priors that may exhibit biased treatment of protected-group predictors.

By Saad Mohammad Abrar, Eesha Kurella, Arnav Dadarya, Naman Awasthi, Kazi Tasnim Zinat, Vanessa Frias-Martinez
arXiv Machine Learning
Sep 24

Dirichlet Process Mixtures of Trees with Gaussian Process Splits: A Bayesian Nonparametric Framework with Posterior Contraction Rate

arXiv:2609. 27930v1 Announce Type: cross Abstract: We propose a Bayesian nonparametric mixture of regression trees with a Dirichlet process prior over tree-parameter pairs, enabling data-driven selection of ensemble size and unifying CART, BART, random forests, and boosting.

By Subhasish Basak, Anik Roy, Sourabh Bhattacharya
arXiv Machine Learning
Jul 10

Bayesian Deep Learning for Discrete Choice

arXiv:2505. 18077v3 Announce Type: replace-cross Abstract: Discrete choice models (DCMs) are used to analyze individual decision-making in contexts such as transportation choices, political elections, and consumer preferences.

By Daniel F. Villarraga, Ricardo A. Daziano
arXiv Machine Learning
Aug 11

Crowd-Sourced Geographies of Income: Using Google Maps Points of Interest as High-Frequency Proxies for Sub-Municipal Income Estimation in Sao Paulo, Brazil

arXiv:2608. 07871v1 Announce Type: cross Abstract: Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade.

By Adrienne C. Kinney, Anya Workman, Ademar Takeo Akabane, Jenna Barac, Paulo Fernando Braga Carvalho, Jeova Farias, Fernando Nascimento, Paulo Ricardo da Silva Oliveira