arXiv Machine Learning By Ravi Satya Durga Prasad Yenugula

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

Read the original on arXiv Machine Learning →

arXiv:2608. 02829v1 Announce Type: new Abstract: Model families train every size from scratch.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 22

Federated Lightweight Fine-Tuning

arXiv:2607. 18343v1 Announce Type: cross Abstract: Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, and gradient compression shrinks it by only a constant factor.

By Radhakrishna Achanta, Will Reed