Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces a per-layer differential privacy (DP) clipping strategy for federated multilingual speech large language models (speech‑LLMs). It demonstrates that standard single‑pool per‑layer DP methods fail due to a cross‑component budget collapse caused by large norm differences between acoustic encoders and language decoders. The authors propose an α‑split two‑pool allocation that normalises encoder and decoder parameters separately, preserving the overall DP guarantee while restoring word error rate performance and providing tighter noise protection for the encoder.
SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.
arXiv:2606. 16461v1 Announce Type: new Abstract: Running large language models locally is often impractical, pushing inference on sensitive text to third-party providers.
arXiv:2407. 08233v3 Announce Type: replace Abstract: Current differentially private learning paradigms face a severe utility bottleneck: DP-SGD degrades performance through noise accumulation over training steps, while aggregation-based approaches such as PATE suffer from data inefficiency due to disjoint data partitioning.
arXiv:2606. 26772v1 Announce Type: new Abstract: Differentially private (DP) training of neural networks is often hindered by the large amount of noise required by gradient-based methods such as DP-SGD, which repeatedly inject high-dimensional noise in parameter space throughout training.
arXiv:2607. 05866v1 Announce Type: cross Abstract: Under a fixed privacy budget, the utility of differentially private (DP) training is ultimately determined by its optimization efficiency.