RW-LoRA introduces a random‑walk approach to fine‑tune LoRA models in a decentralized setting, using a single model token that moves through the network and updates locally. This eliminates the need for global synchronization and reduces communication and computation costs compared to centralized or gossip‑based methods. The authors provide convergence guarantees for non‑convex objectives and demonstrate competitive performance on NLP tasks across various graph topologies.
By Xingran Chen, Rohit Bhagat, Ghadir Ayache, Rawad Bitar, Yanmin Gong, Salim El Rouayheb
arXiv:2510. 13537v2 Announce Type: replace-cross Abstract: On-device deployment of Large Language Models (LLMs) frequently leverages Low-Rank Adapters (LoRAs) to support diverse downstream tasks under tight resource constraints.
By Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli
arXiv:2508. 02932v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has gained popularity as a fine-tuning approach for Large Language Models (LLMs) due to its low resource requirements and good performance.
By Minghao Yan, Zhuang Wang, Zhen Jia, Shivaram Venkataraman, Yida Wang
arXiv:2607. 17181v1 Announce Type: cross Abstract: Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request independently.
By Utopia Meng, Unicornt Zhao, Derek Li, Goalen Gao, Frank Du
Cohere, OpenAI, and AI21 Labs have developed a preliminary set of best practices applicable to any organization developing or deploying large language models.
arXiv:2609.06072v1 Announce Type: cross
Abstract: Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-w...
By Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong
arXiv:2505.14468v2 Announce Type: replace
Abstract: Multi-LoRA (Low-Rank Adaptation) serving allows many specialized LLM variants to share the same base model by attaching lightweight adapters. This...
By Yifan Sui, Hao Wang, Hanfei Yu, Kaiqiang Xu, Yitao Hu, Chen Chen, Jianxun Li, Kai Chen
arXiv:2607. 02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window.
By Minjie Hua, Ning Wang, Peijun Yang, Kai Wang, Shiguo Lian