arXiv:2602. 22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through natural language interaction and tool-mediated decision making.
By Zhiheng Song, Jingshuai Zhang, Chuan Qin, Chao Wang, Chao Chen, Longfei Xu, Kaikui Liu, Xiangxiang Chu, Hengshu Zhu
SWRouter is a new routing method for multi‑turn large language model conversations that uses a similarity‑based context segmentation mechanism to construct prompts and a dual‑metric evaluation framework to separate construction accuracy from router performance. The approach addresses two key challenges in multi‑turn dialogue: preventing information loss or confusion during context construction and evaluating routing quality independently of prompt quality. Experiments on multi‑turn dialogue benchmarks show that SWRouter outperforms strong baselines, improving evaluation accuracy by 16.26% over the best individual large language model and by 8.22% over the Conv‑ID Context baseline.
By Yu Wang, Yuchen Li, Rui Kong, Xinran Chen, Jiamin Chen, Hengyi Cai, Shuaiqiang Wang, Jiashu Zhao, Yulun Zhang, Zhonghao Lyu, Haoyi Xiong, Linghe Kong, Jimmy Xiangji Huang, Dawei Yin
arXiv:2603.04191v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are increasingly serving as personal assistants, where users may share individual preferences over extended interactio...
By Qianyun Guo, Yibo Li, Yue Liu, Bryan Hooi
arXiv:2602. 12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-specific preferences and latent constraints of individual users.
By Yuchen Ma, Yue Huang, Wenjie Wang, Xiaonan Luo, Xiangliang Zhang, Stefan Feuerriegel
The paper evaluates how Retrieval-Augmented Generation (RAG) can improve Large Language Model (LLM) predictions of travel mode choice. Four retrieval strategies—basic RAG, balanced retrieval, cross‑encoder re‑ranking, and a combination of balanced retrieval with cross‑encoder—are tested on three LLMs (GPT‑4o, o4‑mini, o3) using 2023 Puget Sound travel survey data. Results show that RAG boosts accuracy across models, with GPT‑4o plus balanced retrieval and cross‑encoder achieving 80.8% accuracy, surpassing traditional statistical and machine learning baselines and demonstrating strong zero‑shot transfer.
By Yiming Xu, Junfeng Jiao
arXiv:2606. 06178v1 Announce Type: new Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense.
By Jiahao Zeng, Ming Tang, Ningning Ding
arXiv:2606. 18774v1 Announce Type: new Abstract: We present RouteJudge, an online pairwise preference evaluation framework for LLM routing systems, with a public platform available at https://routejudge.
By Guannan Lai, Haoran Hu, Han-Jia Ye
GMTRouter is a personalized large language model router that represents multi‑turn user‑LLM interactions as a heterogeneous graph with five node types—user, LLM, query, response, and turn—to preserve relational structure. Using a lightweight inductive graph learning framework and a user‑conditioned graph sampling mechanism, it captures user preferences from few‑shot data, enabling effective personalization without extensive fine‑tuning. Experiments show GMTRouter outperforms strong baselines, improving accuracy by up to 0.108 and AUC by 0.124, and adapts to new users with minimal data.
By Yihang Sun, Encheng Xie, Tao Feng, Jiaxuan You
arXiv:2607. 10651v1 Announce Type: new Abstract: In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user preferences, and constraints such as opening hours.
By Rongbo Qi, Yaqi Zhang, Shijun Yan, Xuemeng Liu, Xiangrui Cai, Chunyao Song
The paper introduces VIBE‑Bench, a new benchmark designed to test personalized large language models (PLLMs) in a regime where user profile cues and query‑specific preferences do not share the same conceptual space, a situation termed profile‑preference conceptual misalignment (PRCM). VIBE‑Bench contains 3,504 personas, 12,239 dialogues, and a manually verified gold test set, and includes two psychology‑grounded tasks that require cross‑concept preference reasoning beyond surface semantic overlap. Experiments show that existing PLLMs largely depend on shallow semantic correlations and struggle to learn robust cross‑concept mappings, highlighting PRCM as a distinct failure mode for personalization models.
By Yiwen Jiang, Yang Deng, Stephanie Fong, Zimu Wang, Yaling Shen, Wei Feng, Hongxi Yang, Xiangyu Zhao, Zhongxing Xu, Deval Mehta, Xuelian Cheng, Zongyuan Ge
arXiv:2511. 09373v2 Announce Type: replace-cross Abstract: LLMs now tackle a wide range of software-related tasks, yet we show that their performance varies markedly both across and within these tasks.
By Adam \v{S}torek, Vikas Upadhyay, Marianne Menglin Liu, Daniel W. Peterson, Anshul Mittal, Sujeeth Bharadwaj, Fahad Shah, Sujith Ravi, Dan Roth
arXiv:2603.04445v3 Announce Type: replace-cross
Abstract: The rapid growth of large language models (LLMs) with diverse capabilities, costs, and domains has created a critical need for intelligent mo...
By Yasmin Moslem, John D. Kelleher