Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator
Related stories
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Accelerating Protein Language Model ProtST on Intel Gaudi 2
Very Large Language Models and How to Evaluate Them
Evaluating large language models trained on code
Introducing The World's Largest Open Multilingual Language Model: BLOOM
Matryoshka Language Model Suites
arXiv:2608. 09703v1 Announce Type: new Abstract: Training a language model suite classically requires training each model separately and serving them independently.
Red-Teaming Large Language Models
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale.
Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models
Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
arXiv:2601. 21003v3 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration.
Gumbel Distillation for Parallel Text Generation
arXiv:2603. 22216v2 Announce Type: replace-cross Abstract: The slow, sequential nature of autoregressive (AR) language models has driven the adoption of parallel decoding methods.
