arXiv:2509.09710v3 Announce Type: replace-cross
Abstract: This study introduces a Large Language Model (LLM) scheme for generating key attributes of travel diaries in agent-based transportation model...
By Sepehr Golrokh Amin, Devin Rhoads, Fatemeh Fakhrmoosavi, Nicholas E. Lownes, John N. Ivan
The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.
By Yiming Xu, Junfeng Jiao
arXiv:2606. 12657v1 Announce Type: new Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generation.
By Siyu Li, Toan Tran, Lingyi Zhao, Khurram Shafique, Li Xiong
arXiv:2608. 15867v1 Announce Type: cross Abstract: Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging.
By Farbod Abbasi, Zachary Patterson, Bilal Farooq
The paper demonstrates that fine‑tuning large language models (LLMs) on local tourist trajectory data can predict visitor movements under varying conditions. Using 566 trajectories from Wakayama Castle Park, Japan, the authors fine‑tuned Llama‑3.1‑8B, achieving 49.1% accuracy for next point‑of‑interest predictions and maintaining strong performance even on undersampled scenarios such as rainy days. This shows that LLMs can serve as high‑fidelity, context‑aware behavior models for tourist prediction and enable counterfactual analysis of mobility interventions.
By Tatsuya Amano, Hirozumi Yamaguchi
Neural-Bayesian Structure Learning (Neural-BSL) integrates differentiable structure learning with random-utility discrete choice estimation in a single differentiable framework. It keeps observed choices outside the graph to avoid distortion, learns attribute interactions through a structure-weighted network, and propagates interventions by updating attributes in topological order before recomputing utilities and choice probabilities. Evaluations on Seoul stated-preference and London revealed-preference data show Neural-BSL matches conventional benchmarks in predictive performance while uncovering behaviorally coherent dependency structures and revealing downstream traveler and trip adjustments under policy scenarios.
By Hyunsoo Yun, Eun Hak Lee, Jiaru Zhang, Ziran Wang, Eui-Jin Kim
Conventional discrete choice and machine learning models are estimated primarily from observational data and typically treat explanatory covariates as parallel inputs, providing no internal mechanism...
arXiv:2606. 06288v1 Announce Type: cross Abstract: Causal representation learning aims to infer the high-level latent causal concepts that give rise to observed low-level measurements.
By Ankur Garg, Michael Stettler, Aaron Schein, Julius von K\"ugelgen
CALM is a reproducible hybrid framework that combines an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. It executes a closed traveler‑day loop and evaluates each generative module against an empirical, reproducible baseline, using the 2024 New York City Citywide Mobility Survey data. The framework demonstrates significant improvements in mode‑choice accuracy, quantifies trade‑offs through live‑LLM ablation, and supports controlled stress testing and deterministic replay of downstream simulations.
By Yezhou Cheng
The paper investigates how Large Language Models can be used to approximate domain expert priors for Bayesian Networks by extracting probabilistic knowledge about real‑world events. Experiments on eighty publicly available networks across domains such as healthcare and finance show that LLM‑derived conditional probabilities outperform random, uniform, and next‑token baselines. The authors also demonstrate that these LLM‑generated priors can refine data‑driven distributions, especially when data is scarce, and provide the first comprehensive baseline for evaluating LLM performance in probabilistic knowledge extraction.
By Aliakbar Nafar, Kristen Brent Venable, Zijun Cui, Parisa Kordjamshidi
arXiv:2505. 14752v3 Announce Type: replace Abstract: Macro-aligned micro-records are crucial for credible simulations in social science and urban studies.
By Yihong Tang, Menglin Kong, Junlin He, Tong Nie, Wei Ma, Lijun Sun
arXiv:2608.30399v1 Announce Type: cross
Abstract: Large language models (LLMs) exhibit strong semantic reasoning and open-ended generation abilities, but aligning these abilities with structured sequ...
By Yunqi Liu, Yang Zhang, Ruixing Zhang, Liangzhe Han, Yi Qiao, Tongyu Zhu, Leilei Sun