arXiv Machine Learning

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

Thermo-FL is a thermal‑aware federated LoRA fine‑tuning framework that adapts local adapter training and sparse update transmission based on device temperature. It introduces TERRA, a robust aggregation pipeline that uses norm filtering, mask‑aware directional validation, adaptive clipping, and mask‑aware aggregation to defend against Byzantine and communication‑layer adversaries. Experiments on a large‑scale emulator and a Jetson testbed show Thermo-FL improves robustness under adversarial sparse aggregation, stabilizes device temperature, reduces upload size, and preserves utility on GSM8K and BoolQ tasks.

arXiv Machine Learning
5d ago

HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference

HybridInfer is a thermal‑aware reinforcement‑learning router that selects among on‑device, edge, and cloud large language model tiers based on a phone’s thermal headroom and query complexity. Trained offline with a Q‑learning policy, it balances quality, latency, cost, and a thermal penalty, and includes a locality bonus that encourages on‑device execution. In real Android tests on 210 prompts, the learned router outperformed two hand‑tuned heuristics in quality while maintaining the lowest cost, and it proved more reliable and faster than always‑on‑device inference, especially for long queries.

By Simran Koul
arXiv Machine Learning
Jul 7

AdaptiveSD A Stability-Aware, Runtime-Adaptive Speculative Decoding Framework with Multi-Policy Orchestration for CPU-Constrained LLM Inference

arXiv:2607. 03876v1 Announce Type: new Abstract: With the rise of small quantized GGUF-based language models and their increasing use for on-device inference tasks, we have seen the growing need for an approach capable of reliably delivering these models at scale even under severe memory bandwidth constraints such as those imposed by pure CPU implementations.

By Sadra Saremi
arXiv AI
Jun 3

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries

arXiv:2606. 02958v1 Announce Type: cross Abstract: Cross-organization language-model adaptation increasingly faces hard governance constraints: in many deployments, device-level model state-parameters, activations, optimizer state, and per-device updates-cannot be exported outside an administrative boundary.

By Hina Dixit, Punit Kumar, Irene Tenison, Nevasini Sasikumar
arXiv Machine Learning
Aug 5

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

arXiv:2602. 06932v5 Announce Type: replace Abstract: Speculative decoding can significantly accelerate LLM serving, yet most deployments today disentangle speculator training from serving, treating speculator training as a standalone offline modeling problem.

By Junxiong Wang, Fengxiang Bie, Jisen Li, Zhongzhu Zhou, Zelei Shao, Yubo Wang, Yinghui Liu, Qingyang Wu, Avner May, Sri Yanamandra, Ce Zhang, Tri Dao, Percy Liang, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu
arXiv Machine Learning
1d ago

FedSAP: Federated Learning with Structured Adaptive Partitioning for Multi-Domain Heterogeneous Edge Devices

FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.

By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li