arXiv AI

Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes

arXiv:2608. 15867v1 Announce Type: cross Abstract: Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging.

arXiv Machine Learning
Sep 22

CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation

CALM is a reproducible hybrid framework that combines an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. It executes a closed traveler‑day loop and evaluates each generative module against an empirical, reproducible baseline, using the 2024 New York City Citywide Mobility Survey data. The framework demonstrates significant improvements in mode‑choice accuracy, quantifies trade‑offs through live‑LLM ablation, and supports controlled stress testing and deterministic replay of downstream simulations.

By Yezhou Cheng
arXiv AI
Jun 26

Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes

arXiv:2509. 09960v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-quality tabular data is insufficient.

By Mingxuan Jiang, Keyang Chen, Yongxin Wang, Yongsheng Zhao, Ziyue Dai, Yicun Liu, Zeping Li, Qiuyang Zhang, Hongyi Nie, Hongbin Zhu, Sen Liu, Guangnan Ye, Hongfeng Chai
arXiv Machine Learning
1d ago

Relative Transitions, Not Absolute Destinations: A Transfer-and-Ground Framework for Target-Trajectory-Free Human Mobility Generation

The paper introduces Nomad, a transfer-and-ground framework for generating human mobility trajectories without target-city trajectory data. It learns relative transitions from source cities using POI attributes and then grounds these transitions onto a target city’s POI map via a behavior graph and exploration–return walk. Experiments across ten cities show Nomad improves trajectory fidelity and downstream utility by roughly 15% and 3% respectively over adaptation baselines.

By Yidi Wang, Yunhe Zhang, Bangchao Deng, Dingqi Yang, Pengyang Wang
arXiv AI
Aug 25

Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction

The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.

By Yiming Xu, Junfeng Jiao
arXiv Machine Learning
Sep 22

When Does Adversarial Refinement Help? A Negative Result and Open Problem in Adapting R3GAN to Time Series Imputation

The paper investigates whether the stable GAN architecture R3GAN can improve time‑series imputation when adapted to 1‑D temporal data. Using a coarse‑to‑fine refinement framework and a frequency‑domain discriminator, the authors evaluate 14 saved configurations across three datasets and find a negative result: most configurations either show negligible improvement or degrade performance compared to baseline methods. The study highlights that the usual argument—GANs optimize distributional objectives rather than point‑wise ones—does not fully explain the lack of benefit, and it poses an open problem regarding why a learned discriminator fails to provide useful refinement gradients while diffusion denoisers succeed, offering practical guidance on when adversarial refinement may be worthwhile.

By Yufeng He
arXiv AI
Sep 3

UTP-Bench: Uncertainty-aware Travel Planning Benchmark

UTP-Bench is a new benchmark for uncertainty-aware travel planning that evaluates large language models on their ability to generate robust itineraries under real-world stochastic conditions. The dataset covers 504 Indian cities, incorporating attractions, restaurants, accommodations, and multi-modal transportation networks, and includes empirical delay distributions and crowd-density patterns to simulate realistic disruptions. Three new metrics—Buffer Adequacy Score, Crowd-Aware Timing Score, and Transport Delay Absorption Score—measure how well generated plans maintain robustness against transit delays and crowd variability, revealing significant gaps between state-of-the-art LLMs and human-authored itineraries.

By Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana