arXiv Computation and Language

SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations

SWRouter is a new routing method for multi‑turn large language model conversations that uses a similarity‑based context segmentation mechanism to construct prompts and a dual‑metric evaluation framework to separate construction accuracy from router performance. The approach addresses two key challenges in multi‑turn dialogue: preventing information loss or confusion during context construction and evaluating routing quality independently of prompt quality. Experiments on multi‑turn dialogue benchmarks show that SWRouter outperforms strong baselines, improving evaluation accuracy by 16.26% over the best individual large language model and by 8.22% over the Conv‑ID Context baseline.

arXiv Machine Learning
Aug 28

A Survey of LLM Prompt Datasets: Taxonomy, Linguistic Patterns, and Practical Uses

The paper presents a survey of 129 public large language model (LLM) prompt datasets, totaling over 1.22 TB and 673 million instances, and introduces a unified taxonomy for them. By analyzing seven datasets in depth, the authors identify lexical, syntactic, and semantic patterns that differentiate prompts from general text, and evaluate these patterns for tasks such as prompt filtering, source domain routing, and response quality assessment. They demonstrate that a 63‑dimensional linguistic feature set extracted on a CPU can match over 91 % of the F1 score of GPU‑based sentence embeddings while halving latency, and that structural features can effectively route prompts across datasets, though they may negatively impact response quality when prompt length is controlled.

By Yuanming Zhang, Yan Lin, Arijit Khan, Huaiyu Wan
arXiv Computation and Language
Aug 27

VDAR-Router: Adaptive LLMs Routing via Verbalized Query Difficulty Analysis Retrieval

The paper introduces VDAR-Router, a routing framework for large language models that uses verbalized query difficulty analysis to guide model selection. It first generates an explicit difficulty profile for each query, retrieves historical examples with similar profiles, and then estimates model suitability to choose a model based on a reward function balancing performance and cost. Experiments on three datasets show that VDAR-Router consistently outperforms existing baselines in cost‑performance trade‑offs, and case studies confirm that explicit difficulty analysis improves example relevance and routing reliability.

By Yu-Chien Tang, Jun-Chen Hung, Wen-Chih Peng, An-Zi Yen
arXiv Machine Learning
Sep 3

GMTRouter: Personalized LLM Router over Multi-turn User Interactions

GMTRouter is a personalized large language model router that represents multi‑turn user‑LLM interactions as a heterogeneous graph with five node types—user, LLM, query, response, and turn—to preserve relational structure. Using a lightweight inductive graph learning framework and a user‑conditioned graph sampling mechanism, it captures user preferences from few‑shot data, enabling effective personalization without extensive fine‑tuning. Experiments show GMTRouter outperforms strong baselines, improving accuracy by up to 0.108 and AUC by 0.124, and adapts to new users with minimal data.

By Yihang Sun, Encheng Xie, Tao Feng, Jiaxuan You
Hugging Face Trending Papers
Jun 10

Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

Modern conversational agents condition on an ever-growing dialogue history at each turn, incurring redundant attention and encoding costs that grow with conversation length. Naive truncation or summarization degrades fidelity, while existing context compressors lack cross-turn memory sharing or revision, causing information loss and compounding errors in long dialogues.

arXiv AI
Jul 14

Agentic Routing: The Harness-Native Data Flywheel

arXiv:2607. 11399v1 Announce Type: cross Abstract: Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification.

By Xinchen Liu, Hang Zhou, Yingjie Zong, Yuchuan Tian, Liuyang Song, Shuo Zhang, Yulong Li, Wei He, Mengyu Zheng, Runke Liu, Siyang Cheng, Xiang Kuang, Hailin Hu, Kai Han, Yunhe Wang