The paper introduces Poodle, a prototype for just‑in‑time model replacement (JITR) that automatically swaps a large language model with a cheaper, task‑specific model when a recurring task is detected. Poodle reduces inference time by up to 7.5× and saves over $2,200 per 1 M requests compared to a flagship hosted LLM, while maintaining competitive accuracy. The authors argue that model search and transfer learning are essential for efficiently identifying and fine‑tuning these custom models.
By Nils Strassenburg, Boris Glavic, Tilmann Rabl
arXiv:2602. 12323v2 Announce Type: replace Abstract: The widespread availability of fine-tuned LoRA modules for open pre-trained models has led to an interest in methods that can adaptively merge LoRAs to improve performance.
By Haokun Liu, Gyung Hyun Je, Marco Ciccone, Zhenlin Xu, Prasanth YSS, Colin Raffel
UpgradeBench is a decision‑centric longitudinal benchmark that evaluates how fine‑tuned language‑model specialists should be handled when new base‑model releases occur. It covers four consecutive Qwen releases, a continuation checkpoint, six tasks, two model sizes, and OLMo checkpoints with known training lineage, and examines whether retraining, adapter transfer, or other recovery strategies improve specialist performance. The benchmark reveals that upgrade gains vary by task and release interval, that direct adapter copying is sensitive to pretraining distance, and that teacher relabeling can recover specialists without new annotations.
"whyItMatters":"The study provides actionable insights into the cost‑effective management of specialist models across model releases, showing how to balance retraining effort with performance gains."
By Ye Chen, Weining Zhang
arXiv:2602. 05988v2 Announce Type: replace Abstract: Pre-training Large Language Models (LLMs) on web-scale datasets becomes fundamental for advancing general-purpose AI.
By Keith Ando Ogawa, Bruno Lopes Yamamoto, Lucas Lauton de Alcantara, Lucas Pellicer, Rosimeire Pereira Costa, Edson Bollis, Anna Helena Reali Costa, Artur Jordao
The paper introduces FedTAR, a task-aware federated fine‑tuning approach for Mixture‑of‑Experts (MoE) large language models. FedTAR links local client updates to task preferences using routing outputs and Singular Value Decomposition to extract low‑dimensional task coordinates and update directions. It then aggregates updates within and across task clusters, reconstructing the final update to preserve expert specialization and reduce interference, achieving state‑of‑the‑art performance on four benchmark tasks under non‑IID settings.
By Tingqi Wang, Hongyu Ke, Haoxin Wang, Rafal Angryk, Zhipeng Cai
The study evaluates whether a single large language model (LLM) can handle multiple customer‑support tasks or if separate specialist models are preferable. Using 13 models from five families and 200+ checkpoints across eight datasets, the authors find that multi‑task full fine‑tuning consistently outperforms other strategies. They also show that sequential LoRA and model merging can preserve earlier skills and improve off‑task robustness, offering practical guidelines for real‑world deployment.
By Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TN