arXiv AI

ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing

arXiv:2606. 15315v1 Announce Type: new Abstract: Personalized public transit routing in public transit systems remains challenging due to the difficulty of capturing and integrating diverse user preferences into routing algorithms.

arXiv AI
Jun 11

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

arXiv:2602. 22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through natural language interaction and tool-mediated decision making.

By Zhiheng Song, Jingshuai Zhang, Chuan Qin, Chao Wang, Chao Chen, Longfei Xu, Kaikui Liu, Xiangxiang Chu, Hengshu Zhu
arXiv Computation and Language
Sep 11

SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

SWRouter is a new routing method for multi‑turn large language model conversations that uses a similarity‑based context segmentation mechanism to construct prompts and a dual‑metric evaluation framework to separate construction accuracy from router performance. The approach addresses two key challenges in multi‑turn dialogue: preventing information loss or confusion during context construction and evaluating routing quality independently of prompt quality. Experiments on multi‑turn dialogue benchmarks show that SWRouter outperforms strong baselines, improving evaluation accuracy by 16.26% over the best individual large language model and by 8.22% over the Conv‑ID Context baseline.

By Yu Wang, Yuchen Li, Rui Kong, Xinran Chen, Jiamin Chen, Hengyi Cai, Shuaiqiang Wang, Jiashu Zhao, Yulun Zhang, Zhonghao Lyu, Haoyi Xiong, Linghe Kong, Jimmy Xiangji Huang, Dawei Yin
arXiv Machine Learning
Jun 30

Synthetic Interaction Data for Scalable Personalization in Large Language Models

arXiv:2602. 12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-specific preferences and latent constraints of individual users.

By Yuchen Ma, Yue Huang, Wenjie Wang, Xiaonan Luo, Xiangliang Zhang, Stefan Feuerriegel
arXiv AI
Aug 25

Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction

The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.

By Yiming Xu, Junfeng Jiao
arXiv Machine Learning
Sep 3

GMTRouter: Personalized LLM Router over Multi-turn User Interactions

GMTRouter is a personalized large language model router that represents multi‑turn user‑LLM interactions as a heterogeneous graph with five node types—user, LLM, query, response, and turn—to preserve relational structure. Using a lightweight inductive graph learning framework and a user‑conditioned graph sampling mechanism, it captures user preferences from few‑shot data, enabling effective personalization without extensive fine‑tuning. Experiments show GMTRouter outperforms strong baselines, improving accuracy by up to 0.108 and AUC by 0.124, and adapts to new users with minimal data.

By Yihang Sun, Encheng Xie, Tao Feng, Jiaxuan You
arXiv AI
Sep 2

VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences

The paper introduces VIBE‑Bench, a new benchmark designed to test personalized large language models (PLLMs) in a regime where user profile cues and query‑specific preferences do not share the same conceptual space, a situation termed profile‑preference conceptual misalignment (PRCM). VIBE‑Bench contains 3,504 personas, 12,239 dialogues, and a manually verified gold test set, and includes two psychology‑grounded tasks that require cross‑concept preference reasoning beyond surface semantic overlap. Experiments show that existing PLLMs largely depend on shallow semantic correlations and struggle to learn robust cross‑concept mappings, highlighting PRCM as a distinct failure mode for personalization models.

By Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge