arXiv AI By Yi Li, Chen Li, Jiexiong Liu

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

Read the original on arXiv AI →

arXiv:2607. 13093v1 Announce Type: cross Abstract: On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.