arXiv Machine Learning By Shiva Shrestha, Kazi Shaharair Sharif, Zongxing Xie, Jiajing Huang, Anhao Xiang, Honghui Xu

Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AI

Read the original on arXiv Machine Learning →

Thermo-FL is a thermal‑aware federated LoRA fine‑tuning framework that adapts local adapter training and sparse update transmission based on device temperature. It introduces TERRA, a robust aggregation pipeline that uses norm filtering, mask‑aware directional validation, adaptive clipping, and mask‑aware aggregation to defend against Byzantine and communication‑layer adversaries. Experiments on a large‑scale emulator and a Jetson testbed show Thermo-FL improves robustness under adversarial sparse aggregation, stabilizes device temperature, reduces upload size, and preserves utility on GSM8K and BoolQ tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference

HybridInfer is a thermal‑aware reinforcement‑learning router that selects among on‑device, edge, and cloud large language model tiers based on a phone’s thermal headroom and query complexity. Trained offline with a Q‑learning policy, it balances quality, latency, cost, and a thermal penalty, and includes a locality bonus that encourages on‑device execution. In real Android tests on 210 prompts, the learned router outperformed two hand‑tuned heuristics in quality while maintaining the lowest cost, and it proved more reliable and faster than always‑on‑device inference, especially for long queries.

By Simran Koul
arXiv Machine Learning
Jul 7

AdaptiveSD A Stability-Aware, Runtime-Adaptive Speculative Decoding Framework with Multi-Policy Orchestration for CPU-Constrained LLM Inference

arXiv:2607. 03876v1 Announce Type: new Abstract: With the rise of small quantized GGUF-based language models and their increasing use for on-device inference tasks, we have seen the growing need for an approach capable of reliably delivering these models at scale even under severe memory bandwidth constraints such as those imposed by pure CPU implementations.

By Sadra Saremi
arXiv AI
Jun 3

Echelon: Auditable Aggregate-Only Language-Model Adaptation Across Privacy Boundaries

arXiv:2606. 02958v1 Announce Type: cross Abstract: Cross-organization language-model adaptation increasingly faces hard governance constraints: in many deployments, device-level model state-parameters, activations, optimizer state, and per-device updates-cannot be exported outside an administrative boundary.

By Hina Dixit, Punit Kumar, Irene Tenison, Nevasini Sasikumar