arXiv Machine Learning By Ligong Han, Kai Xu, Hao Wang, Ruijiang Gao, Akash Srivastava

Layer-Parallel Inference Reduces Encrypted Nonlinear Depth in Transformers

Read the original on arXiv Machine Learning →

arXiv:2607. 04819v1 Announce Type: new Abstract: Fully homomorphic encryption (FHE) enables computation on encrypted data, but practical encrypted Transformer inference is bottlenecked by the sequential composition of many nonlinear blocks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 3

HEAT: Faster Fully Homomorphic Inference via Approximations-Weights Co-Adaptation

HEAT introduces a fine‑tuning method that treats the number of iterations used to approximate nonlinearities in fully homomorphic encryption (FHE) as learnable parameters, allowing them to co‑adapt with model weights. By optimizing iteration counts per nonlinearity, HEAT reduces the required iterations, bootstraps, and overall latency for encrypted GPT‑2 decoding while improving decode agreement. The approach achieves a 3.1× reduction in iterations, a 1.6× reduction in bootstraps, and a 1.4× speed‑up in end‑to‑end latency without changing the model architecture or requiring retraining from scratch.

By Alessandro Zirilli, Davide Marincione, Evgenios M. Kornaropoulos, Giuseppe Ateniese, Emanuele Rodol\`a
arXiv Machine Learning
Jun 3

Why Are Linear RNNs More Parallelizable?

arXiv:2603. 03612v3 Announce Type: replace Abstract: The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability.

By William Merrill, Hongjian Jiang, Yanhong Li, Anthony Lin, Ashish Sabharwal