Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
arXiv:2607. 13093v1 Announce Type: cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy.
mmFHE is the first system that runs the entire cloud-side mmWave sensing pipeline—including DSP and machine‑learning inference—under fully homomorphic encryption. It encrypts range profiles on an edge device, then processes them homomorphically on a semi‑honest cloud using a library of seven data‑oblivious FHE kernels that replace standard DSP routines. The authors demonstrate the approach on vital‑sign monitoring and gesture recognition, proving input privacy and data obliviousness, and show negligible accuracy loss (84.5% vs. 84.7%) with practical GPU latencies on commodity hardware.
arXiv:2607. 13093v1 Announce Type: cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy.
arXiv:2606. 16359v1 Announce Type: cross Abstract: Fully Homomorphic Encryption (FHE) enables privacy-preserving machine learning but incurs extreme computational and memory overhead.
The paper introduces GASHE, a gradient‑aware selective homomorphic encryption scheme that encrypts only those gradient components exceeding a differential‑privacy‑calibrated sensitivity threshold, rather than encrypting all parameters. Building on GASHE, SecureDrive‑FL combines DP‑SGD with this selective encryption to form a closed‑loop DP+HE privacy pipeline for federated driver monitoring. Experiments on a ten‑class distracted driver classification task show that SecureDrive‑FL retains the poisoning resistance of DP‑SGD while also defending against Man‑in‑the‑Middle interception, with only an 8–10% runtime overhead.
arXiv:2609. 01945v1 Announce Type: cross Abstract: Federated Learning enables multiple clients to train a shared model while keeping their local datasets isolated.
HEAT introduces a fine‑tuning method that treats the number of iterations used to approximate nonlinearities in fully homomorphic encryption (FHE) as learnable parameters, allowing them to co‑adapt with model weights. By optimizing iteration counts per nonlinearity, HEAT reduces the required iterations, bootstraps, and overall latency for encrypted GPT‑2 decoding while improving decode agreement. The approach achieves a 3.1× reduction in iterations, a 1.6× reduction in bootstraps, and a 1.4× speed‑up in end‑to‑end latency without changing the model architecture or requiring retraining from scratch.
SecureDrive‑FL combines differential privacy (DP‑SGD) with a novel Gradient‑Aware Selective Homomorphic Encryption (GASHE) scheme to protect federated driver‑monitoring models. GASHE encrypts only gradient components that exceed a DP‑calibrated sensitivity threshold, avoiding full‑parameter encryption. In experiments on a ten‑class distracted driver task, SecureDrive‑FL matches DP‑SGD’s poisoning resistance while also defending against Man‑in‑the‑Middle attacks, adding only 8–10% runtime overhead.
arXiv:2609.16898v1 Announce Type: cross Abstract: Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with...
SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.
arXiv:2607. 20890v1 Announce Type: new Abstract: On-device federated learning (FL) enables privacy-preserving and personalized model training on resource-constrained devices such as smartphones and IoT nodes.
Fully Homomorphic Encryption (FHE) enables computations to be performed directly on encrypted data while preserving data confidentiality. However, its practical applications remain limited by high computational costs and development complexity.
arXiv:2607. 29221v1 Announce Type: cross Abstract: We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untrusted but powerful server, in the setting where the client holds both the input and the model, and the server must learn neither.
arXiv:2606. 00279v1 Announce Type: cross Abstract: Verifying claims about AI workloads is a pre- requisite for credible AI governance of covert adversaries (who comply with monitoring only when detection likelihood is high), yet the ap- parent non-determinism of GPU floating-point arithmetic forces auditors to accept approximate output matches.