← Back to all news
Hugging Face Blog September 14, 2021

Introducing Optimum: The Optimization Toolkit for Transformers at Scale

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Blog
Jul 8

Native-speed vLLM transformers modeling backend

llms
More like this →
Hugging Face Blog
Sep 14, 2021

Hugging Face and Graphcore partner for IPU-optimized Transformers

llms
More like this →
arXiv Machine Learning
Jun 16

LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers

arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.

By Abhishek Shukla, Anikeit Khanna, Ankur Sinha, Faiz Hamid
llmsfine-tuning
More like this →
arXiv AI
Aug 10

Stability of Transformers under Layer Normalization

arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.

By Kelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai, Stanley Osher, Krishna Kumar, Markos A. Katsoulakis
llms
More like this →
Hugging Face Blog
Aug 17, 2022

A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes

llms
More like this →
Hugging Face Blog
Feb 26

Mixture of Experts (MoEs) in Transformers

llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e